As we continue developing our software, we accumulate a growing amount of technical debt just to keep the system running. But I believe we are on the brink of an even larger issue. Cognitive debt.

Hope you enjoy this reading, all feedback is welcome.

  • MagicShel@lemmy.zip
    link
    fedilink
    English
    arrow-up
    1
    ·
    3 days ago

    If you weren’t using LLM for the message, this is on me.

    No worries. These days if you haven’t made a false LLM accusation, you aren’t looking hard enough.

    Often the projects that other people tried to make something new and fumbled so hard that the company owner had to buy my labour and knowledge to fix the shit.

    Oh we work on the same kinds of projects for sure! We have several legacy microservices we are required to support.

    So I think you may not have a full cycle on the software development, the green field is always easy, and this wasn’t never the problem.

    Well firstly, you say greenfield was never the problem, which isn’t the conclusion one comes two when your three examples are all failed greenfield projects. That said…

    Yeah I conflated a number of things which could have used better clarification. We have a harness for doing AI-led greenfield development. And it is unproven, and to be frank I’m skeptical as fuck about it. I think it’s a terrible idea. We also have a more modest process for doing human-led development with AI assistance, and it requires some up front work writing ai-facing documentation, but it does improve delivery speed.

    That being says, my production support tools, which were written by AI and are orchestrated by AI are a massive productivity boost. I can run 5 investigations in parallel and get them all done in less time than it would take me to do one. Make no mistake, Claude’s interpretation of the data is frequently awful, but it pulls all the raw data from all the relevant systems, and with a little expert guidance to challenge incorrect analysis, I can confirm and refute hypotheses faster than I could manually copy a trace id from Kibana and search it in Dynatrace. It’s by far the biggest productivity gain.

    However I think the process of requiring everything to be documented from the architecture to the acceptance criteria to the API is a good idea and enables AI to make meaningful contributions. This harness is not going to work for legacy maintenance because legacy will never be documented as well as a project which has never allowed code without documentation from every angle.

    Not having the full context of other services and without the good test case for it, the product will be fated to fail.

    We have have swagger and, in many cases, code, and we run into the same issues with standard development — we have to reach out to other teams to know how to create and remove test cases for automated integration tests. But microservices are pretty well documented for API and usage. This isn’t the problem you think.

    So I would suggest you to validate your own assumptions on the topic.

    I am, as we speak. We are running parallel development efforts on the same design. And between you and me, I expect the human-driven development to be better. That said, AI thus far is no less annoying about discovering every little question that isn’t specified or documented. I don’t like the process and I don’t have faith in the process, but so far the results are positive.

    That’s very funny. The same academic research who found out that the usage of LLM delayed the deliverables in ~19% had a section where the developers using the tool thought (wrongly) that they were ~25% faster.
    You are only proving the paper.

    I’m aware of the research. And when I was tasked to increase AI adoption the first thing I did was explain this paper to my boss and tell him we need REAL metrics. Actual story points delivered over time per team. And I look for every flaw I can find to tear down those number’s and arrive at some truthful value. If you spent 3 story points doing tasks only required because of AI, like documenting or writing skills instead of code, that’s not productivity at all.

    And here’s what I’ve found in delivered story points: at first AI lowered real productivity because while story points went up, they were being spent on tasks that were unnecessary under traditional development. As you say, the self-reported numbers (all we had at first because it takes time to see actual trends) all showed more productivity, but it was a lie when you look back at actual delivery.

    But over time, we spend less time on ai-driven tasks and actual delivered work has increased. Reports show higher gains for QA than for for development.

    No one is a bigger skeptic about these numbers than me. While I was hired over other candidates specifically for my AI knowledge, I’m still a technical lead first, and responsible for the quality of code we are delivering. I test everything AI writes, I use deterministic processes everywhere I can, and I continually push back. I tease apart every response to find flawed reasoning.

    As a result, I’m not showing 10x gains companies are hoping for that suddenly evaporate when they hit the real world. I’m showing 20-30% actual gains across 8 devs and 4 QA because we are enabling devs rather than replacing them.

    Which is why I’m so skeptical about this dev harness. I think it’s likely to be wasted time on an overengineered solution that works worse than developers. It takes techniques I’m slowly improving and expanding, and leapfrogs straight to the race track, and I think despite being built on a good foundation, it just threw my proven techniques into the AI blender and got out a harness that is going to fail.

    But I have to prove it. The questions AI has raised so far have highlighted many gaps in the initial design, and the code it produces looks good on code review, but I’m fully aware that I don’t have the totality of the code in my mind when I review it because I’ve only seen the code once, on an earlier code review.

    It may turn out that the harness’ greatest value is in improving architecture and project planning by finding all of the ambiguities and inconsistencies in the specs faster than the devs through a prototyping process that delivers an actual working poc according to specs — and then we turn those designs over to devs to implement correctly. The value of that will depend on tokens spent.

    I’m looking for the similar size text, this usually fits in the commute time, have a deeper conversation/meaning.

    Well I’m certainly arguing against my initial preposition here, aren’t I?

    I am enjoying the conversation, though this is probably not the best format for it. But I appreciate your thoughts. It would be nice to get a bit more credit for my experience and ability both with LLMs and development leadership in general, but of course you don’t know me and credibility takes time to achieve. I feel like maybe I’m coming across as defensive when I’m trying to establish my level of expertise. It certainly increases word count, though I don’t know if it does anything of actual value.

    Anyway, it’s good to read other voices that are both skeptical and cautiously optimistic that there is something here of value.