top of page

Not Everything Should Be Level 4 AI

I want to start off by saying this blog post is “level 2 AI”. That means you shouldn't be surprised if a sentence or two feels a bit LLM-esque, but you should also expect that the entire piece has a bit more carbon than silicon in its vibes. I am stating this upfront to follow the practices this blog introduces and let you experience it from a consumer side.


The main focus of this blog is getting the most out of communication. Producing content, often written content, has many uses. Some are for the author, where it helps them reason through ideas, and some are for the reader, where it helps them understand and possibly deliberate on a topic more easily. In either case, writing shouldn't be about the quantity of words you can put on paper, and I want to look at ways that I have addressed this in a time when it feels like that is exactly what LLMs are working towards.


The problem with AI-First drafting

Having AI draft an initial pass at content can deteriorate the value in a few different ways that only really become clear once you dig into actually putting the writing to use.


As an author, if you sit down with a blank page, it can be hard to get started. So today it is absolutely reasonable to reach for an LLM to move you from creation into editing mode, which is viewed as easier. The challenge is doing this without falling into the trap of anchoring. With so many words written, it is difficult sometimes to evaluate the overall messaging and tone without getting bogged down in specific edits. This means our human brain power isn't actually applied to the highest-value work.


As a reader, it is absolutely infuriating to sit down and realise you are reading AI when you thought you were reading a colleague's thoughts. It is also tiring to slog through AI writing even when you knew it was drafted using AI. I mean, if I am being honest, it doesn't just sneak up on you; it hits you like a wave, and in today's fast-paced, AI-driven world, that shift from quiet unease to undeniable realisation doesn't just change how you read a single document; it fundamentally transforms the entire dynamic of trust between colleagues. At the end of the day, that's a cost none of us can afford to ignore.


See? But promise, that was the last generated sentence in this blog!


A useful AI usage scale I invented half as a joke

What I started to realise was that part of the issue was the writer’s choice to use AI, and part of the issue was the readers’ assumptions about that choice, given they had to deduce it themselves.


In the good old days of late 2025, I used to assume people had spent time writing things and painstakingly providing feedback just to be told: "oh, this was just an AI draft, I will have a look later". Cue absolute rage. But a year on, it still doesn't feel simple to ask. It kinda feels like when your book club friend is yammering on as if they read the book, but it just doesn't line up. It is weird to ask from the outside, even if it does affect your experience as a fellow group member. And as the delinquent reader, it can be embarrassing to say you didn't keep up if the culture isn't friendly to that scenario. Often the fear of accusing and getting it wrong can easily outweigh the value in getting shared understanding, so you just make assumptions and work off those even when that can create its own problems.


To get past that, I decided I wanted to normalise saying what level of AI I had used, hoping others would clock on. I drafted a blog with AI, read it and edited where needed, but overall it definitely still sounded like AI. Not my proudest work, but given the topic and context, it made sense, and I wanted to see if others agreed. So what I did was I posted it to my team with a comment that it was level 3 AI, which I defined as "drafting but manual heavy revision".


And posted this as the levels I loosely referred to:


Announcing what amount of AI was used in my work to set expectations on review.
Announcing what amount of AI was used in my work to set expectations on review.

By framing the context, the resulting conversation focused on the question at hand: is this the right shape and type of content, not on the specifics of AI language used. Of course we needed to fix up the language before publishing, but first I wanted to gain confidence that the content was worth the energy, and this helped focus that conversation.


It isn't a scale of how hard you tried

I have had some learnings since using this and since reflecting on how others do. First of all, I am surprised how much I don't immediately dismiss level 5 AI as much as I thought I would. I also don't assume lower numbers == more effort or better. What I do judge quite quickly is whether I feel like the level done aligns with the level I think that writing should have. For example, if we are designing a new feature, I expect no higher than 3 (and likely level 2). If it is a summary of where we are with certain metrics, I am not surprised to see 5 or maybe 4. I would be equally critical of a teammate who chose to spend their manual time collating data as I would be of a teammate who offloaded the future of our product to something that explicitly aims for "average".


To my surprise, other people started using it

Over a few weeks, I iterated on the initial proposal and ended up with the following scale.


  1. No AI used

  2. Research only, completely human-drafted

  3. Co-drafting and/or copy-edited with heavy manual revision

  4. AI drafted, light manual review/revision

  5. Raw AI output from minimal prompting


Yes, the first thing you will notice is that the engineers engineered and went to a zero base scale because, "as silly as it sounds (but was triggered by your example, “This proposal is AI level 1”), could we start counting from zero?"


You might also notice it didn't really change much from the original, yet this is the version that has now been implemented at Syntasso, with a subgroup of the CNCF Platform Engineering Technical Community Group, and with a number of companies that members of that CNCF community work for. I think it is because it just clicks without needing much explanation.


More than the scale, though, I find the introduction to why it exists is most helpful. I now try to inject the need to always think about the reader/listener/consumer before publishing something. Despite my constant struggle with it, I try my best to implement ideas from The First Minute and realise that getting the best outcomes demands thinking about how to set those up early. This also goes for conference talks where the audience is the main driver for success.


So my most common suggestion as people start creating content is to answer "who is the audience and what do you want them to do or feel after reading this?" For example, with this blog, my audience is software professionals who collaborate with their teammates. And I want them to use communication intentionally rather than as a byproduct.


How I drink my own champagne

So if my goal is always to think of the reader before sharing anything, then I need to be intentional about how my choice of format, length and style affects their ability to take in the content more than what is easiest for me to produce. As a clear example of this, for some audiences it may be most effective to provide more than one option for consumption, which is most certainly not the easiest thing for me to do!


As my lovely and skilled colleague Daniel Bryant loves to say, I bury the lead too much (and yes, the lead for this blog got completely rewritten on edit). This is a weakness of mine personally, but it is also a reality that summarising something often requires hashing out the content first to a level that can be effectively summarised! That is why, in practice, I often find myself writing a post, then going back to the beginning and adding a clear TL;DR line, and reorganising the body to have useful headers and order.

Separately, I find that labelling the AI use makes me rethink whether I am content with my choices.


Sometimes I realise my research was too narrow if I am too low on the scale. Other times I realise I need to do a better proof read/edit before asking someone else to do my work for me. This is very similar to the impact of adding a test plan to a PR. When I am asked to summarise the value of a PR into my product, I have to think about if this is actually user-friendly and helpful. When asked to also provide context on what I have tested, I realise that sometimes I didn't do the level of coverage necessary and can backfill before asking for time and energy from a teammate.


In both of these cases, applying these "AI use scales" isn't actually about AI at all. While we are asking the question "how did you use AI", the real outcome is a review of how well you applied the tools at your disposal and a prompt to be more effective at communication.


Typing makes you reread yourself. Talking usually doesn't.

While not directly related, I think that another tangential point is about how I actually interact with AI when I do use it. And I think that it is another example of how prioritising outcomes needs to outweigh production. I got really into AI-supported speech-to-text for a while. I loved sort of rambling away and seeing what happens. Now I use it much less, but still consider it a fabulous tool in my toolbox.


Sometimes the fidelity of my thinking is just rambles. I need to talk something through just like I would with a teammate or a rubber duck. The act of rambling is actually what will drive out clarity. In these cases, using AI to turn rambling speech into clean, readable text. But I don't think in a structured way when I am speaking out loud, so using speech-to-text to reduce the load on me to produce output can often be solving the wrong problem. It is easing my ability to produce content without materially improving the ability for others to consume that content.


When I type, it is slower (despite a respectable 91 wpm, thanks, type racer!) I know that my brain moves faster. This lets me reread sentences as I go, sometimes moving or editing them or even just realising that they were filler fluff. When I am talking at a screen, the sentences flow out roughly in the order they came to mind, and I don't commonly stop and edit, as that would go against the value of the forum.


None of that makes speaking to an agent the wrong choice, just one that you should do intentionally. Being intentional about how we input can set us up for success in being intentional about the output.


Final thoughts

Agile sticky notes were never about the stickies, but about the crisp and dynamic conversations they drove. If you have to pick between scenario writing and automating, BDD practitioners advocate for scenario writing. The recording part of ADRs is only valuable when you are recording the why as much as the what. These examples prove that the journey is worth as much as the output, and I fear that with the ease of content generation, we are losing that journey. But I equally fear a world where using AI is frowned upon rather than wielded with care. I think bringing the conversation into the light and including it in how we plan and evaluate our work can avoid both extremes.


I hope that this is helpful for you and your team. And in the spirit of being intentional. I hope you take this and evolve it for your context. Focus on the impact, not on following these ideas to the letter.

Comments


bottom of page