OrgLens: 2. Making an organisational health check people can actually use

In my first note on OrgLens I asked whether an AI agent could diagnose an organisation using Stafford Beer’s Viable System Model (VSM). The answer was yes, more or less.
The reports were written for someone like me using terms such as System 3*, algedonic channels, and requisite variety. Those terms can be useful if you know VSM, but most people who might benefit from an organisational health check do not, and shouldn't need to.
So the next question was whether OrgLens could give someone with no systems thinking background a useful report about their organisation. The answer? Yes, with caveats. I had to make plain language and checking part of the software, rather than asking the AI to remember them. I also moved from an agent that chose its own next steps to a fixed sequence of steps.
As before, I made the meta-level decisions, and then Claude Code turned that into instructions and specifications for Raven, an AI agent tool.
Raven wrote and amended code, built report templates, ran checks, and produced trial reports. Most runs used OpenRouter’s free router. I reviewed the results and decided what to test next.
What does 'usable' even mean?
I wanted a person to be able to describe their organisation in ordinary language and receive a report that they could understand, trust, and act on, so we tested six things:
- People can start without VSM terms, forms, or structured data files.
- The report uses plain language and keeps the theory out of sight.
- It says what is working, what may need attention, and what to discuss first.
- Claims can be traced back to something someone said.
- A run finishes with a report rather than an error.
- Plain language does not make the diagnosis less accurate.
I tested this with three fictional organisations. To make things interesting, two were presented as tidy interview notes while a third arrived as a long and apologetic message from a volunteer, with the sort of digressions people make when explaining a real organisation.
To be clear, I used simulated readers: AI agents playing an operations manager, a garden-centre owner sceptical of consultants, and a new trustee. They read reports without preparation, said what they took from them, listed words they did not know, and chose what they would do first.
It's an approximation of the real thing and useful for finding obvious problems. But it's obviously not a substitute for testing with real people.
Building in 'plain language'
The first reports did poorly on clarity with six readers gave the best phase-one reports a score of 2/5. On average, each reader found 13 terms or phrases they did not understand. That's not great.
One said they would skip straight to the recommendations and strengths next time. Another called the diagram code at the end “gibberish”. I found that quite realistic. 😂
So, yes,asking a model to write in plain English worked sometimes, but I wanted the result to be consistent. Consequently, I asked Claude to tell Raven to give each kind of finding a plain-language form. Instead of VSM labels, the report asks questions a manager might ask such as:
- "Is anyone looking ahead?"
- "Does anyone check what really happens?"
- "Can bad news reach the top quickly?"
- "Do today’s pressures crowd out tomorrow’s?"
Report titles now uses the organisation’s own words. For example, "Warm Homes advice has to ask permission for too much" replaces S1-LOW-AUTONOMY,
Raven also created a command that writes the basic report structure before the AI writes the summary and recommendations. This left the model with the parts that need judgement, rather than asking it to write every part from scratch. Horses for courses.
Get that jargon out of my face
The first plain English reports included a short VSM appendix for readers who were curious. However, as with humans, nobody cares, so moving the appendix into a separate file made the main reports much easier to read. Most of the remaining unfamiliar terms were the organisations’ own jargon, quoted from the interviews, such as "DNO" and "MCS certificate".
The lesson is pretty simple and something I should probably apply to my consulting practice: it's not enough to remove jargon from the opening paragraphs. If it appears anywhere in the document, it still shapes what readers think of it.
Checking claims, not just asking nicely
Somewhat unsurprisingly, the first phase showed that asking an AI to check its own evidence often produces a ritual rather than a real check. It's surprising how lazy it can be – just like humans, I guess. 😅
So I asked Claude to instruct Raven to add checks that run every time:
- Every quotation is compared with the source notes.
- The report is checked for contradictions, such as calling safeguarding both "excellent" and a serious weakness.
- The model starts from a template, rather than inventing its own structure
- Small formatting mistakes are repaired before the report is written.
That means the rule is now simple: if the report cannot point to something someone said, it either goes under "things we're less sure about" or it is removed.
These checks caught invented quotes and claims that contradicted the notes, but did not solve every problem. Models still skip steps, and a well-formed report can still rest on a weak model.
Raven was not built for this audience
What was great about Raven is that it could find the appropriate tool from an ordinary request such as "can you give our makerspace a health check?". But the rest of the experience was less friendly for non-technical people.
Raven asks users to approve shell commands, start sub-agents, and approve changes to files. It saved reports in hidden folders and could leave someone watching "1 running" for ten minutes, with no indication of what was happening.
That's why having Claude Code running a frontier model translating and instructing Raven as an orchestrator is particularly useful. After changing the settings and instructions, one browser run turned the Riverside message into a 1,057-word report in 13 minutes. It used no systems-thinking jargon, checked all 19 quotations against the original message, and found all eight planted problems in its model.
While not super-speedy, that's using a free models router and was good enough to show the method could work. However, it was still not a good enough experience to give to a non-technical person.
From agent to web app
At this point, the agent was mostly there to just remember to run a set of tools in the right order. So we didn't need the agent any more and could instead run a fixed pipeline:
- Read the person’s notes and build a structured model of the organisation.
- Repair and validate that model.
- Run the structural checks.
- Produce a plain-language draft.
- Check quotations, contradictions, and missing sections.
- Write the final report.
Only two of these steps need an AI model: building the model from people’s words, and writing the report. Everything else runs in code, so I asked Claude to instruct Raven to build that pipeline.
If the report writer will not correct an invented quote after one retry, the software removes the quote marks. Nothing appears as somebody’s words unless they said it.
Choosing models based on evidence
I purposely had no credit in my OpenRouter account so that Raven didn't 'cheat' and always used free models. I wanted to see how far it could go without me having to burn through hard-earned cash.
The thing is that the free router chooses a model for each call meaning that a single diagnosis could pass through six to eight different models. When I tried "pinning" one free model that actually made things worse: rate limits and failed replies meant that none of four runs completed.
I tested candidate models on all three fictional organisations, and the best free options for building the organisational model were Nemotron 3 Ultra (550B), minLing 3.0, and Nemotron 3 Super (120B). OrgLens uses an ordered list. It tries the best available model first, then falls back to the next one.
For those wondering, I later did try Jev Router, a paid option that chooses a model and a level of reasoning for each request. It was faster and a little more accurate, finding ~86% of planted problems, compared with 83% for the best free model. A health check took 1 to 6 minutes and cost $0.05 - $0.16. But I didn't think that was good enough of an improvement to make it the default. Perhaps I'm just a cheapskate.
What the web app does
I'm not quite ready to realease it to the world yet, but essentially what happens is: the user opens a page, reads a privacy note, writes about their organisation (and/or uploads notes), and starts the OrgLens process.
The page explains the steps involved in plain language. It shows progress, says how many jobs are ahead in the queue, and gives the finished report in the browser. The privacy notice comes first for a reason as what people write is sent to AI models run by other companies. Free models may keep it, that's usually part of the deal. That could matter for non-public organisational information.
I sent all three test cases through the app at once, and the generated reports were rated 4.0 to 4.3 out of 5 for clarity and 4.3 for usefulness. For what it's worth, all nine simulated readers said they would share them.
Limitations
This is obviously still an experiment and there are things I need to figure out when I'm back from travelling. Current caveats:
- The readers were AI agents playing people, not humans.
- The same AI setup helped build the tools, write the fictional organisations, and run the evaluation.
- There were only three readers per condition. Small differences are likely to be noise.
- Free model availability and rate limits change often.
- The tool runs locally, one job at a time. A public service would need sign-in, a privacy policy, and paid models with clear data terms.
The work so far does suggests, though, that a plain-language organisational health check is possible. The most reliable parts are the ones enforced in software: quote checks, section checks, and contradiction checks.
The hardest parts remain human ones: deciding what matters, whether a finding is fair, and what an organisation should do next. Humans and machines working in harmony?
For the sake of transparency, Claude created the first draft of this note. Image: Are.na