OrgLens: 1. Can an AI agent do systems thinking?

Systems thinking is a way of looking at organisations as sets of connected parts rather than as collections of individuals or departments. The Viable System Model, or VSM, is one approach within it. It asks whether an organisation has the parts it needs to keep going: people doing the work, ways of coordinating that work, checks on whether the work is happening, people looking ahead, and a way of agreeing what the organisation is for.
There are many different schools of systems thinking. Very broadly, there's a spectrum between: 'hard' cybernetic-style approaches such as VSM and 'soft' approaches such as Critical Systems Heuristics (CSH). My view, which seem to be common in the modern landscape, is that we should be adopting a multi-methodologies approach. Also, we should also remember that they are merely lenses; there is no capital-t "Truth" to be had here.
So, because I've been experimenting so much with AI over the past few years, I've bee wondering whether today’s AI agent tools could help with the kind of organisational diagnosis I do using systems thinking. In particular, I wanted to test how good AI is applying VSM.
So I spent an evening and a night running an experiment. In essence, I told Claude Code what I wanted to do, and it implemented Raven, a new open source project from EverMind, to do it. Raven is described as "the harness of harnesses" which, in plain terms, means that it's a host agent: an AI system that breaks a task into steps and passes those steps to other, more specialised agents. Raven can remember work between sessions and has an experimental feature for shaping a "digital employee" from your own materials and feedback.
On paper, that sounded INTERESTING and like a good fit for this experiment.
This note covers what was built, how it was tested, and what I learned as a result. The TL;DR is that the approach worked better than I expected on the parts I thought would be hard, and worse on the parts I thought would be easy. C'est la vie.
Raven resembles a viable system
Before getting started, it was interesting to note that Raven’s own design maps quite neatly onto the VSM:
- System 1: Operations – Specialist agents doing the work
- System 2: Coordination – Task graph that puts the work in order
- System 3: Control – Module that checks each action before it runs
- System 3*: Audit – "Oncall" agent, which monitors long-running work
- System 4: Intelligence – Research agents and scheduled monitoring
- System 5: Policy – Persona and "Curator" which set identity and limits
I don't think this resemblance is accidental. After all, anyone building a system of autonomous parts that has to remain coherent ends up solving the problems the VSM is designed for, whether or not they have read the work of Stafford Beer (who came up with it).
Ch-ch-ch-check yourself (before you wreck yourself)
Before the agent got involved, Claude decided to split the work in two:
Modelling – Reading and describing the organisation in VSM terms. This requires judgement, so it is the agent’s job.
Checking – Looking for the structural problems the VSM is designed for: missing coordination, a System 4 that nobody owns, a board spending all its time on operations, or no route for an "alarm" to reach the top. This does not require judgement, so ordinary code does it. The idea is that the same code should give the same answer every time.
The model is a plain YAML file. YAML is just a simple text format for structured information, meaning that every part of the model can carry the evidence behind it, quoted from the source material.
The checker produces findings, and each finding comes with a question to ask the organisation next. It's important to note that these are hypotheses, not verdicts.
This is a good way of splitting things, because:
You can argue with a model file in a way you cannot argue with a confident paragraph
Checks that do not depend on a model’s "mood" can be repeated
It puts the AI where it is most useful: turning messy notes into structure
The method itself went into a "skill": a folder containing a SKILL.md file of instructions, a VSM primer, and a model template. Skills are a common format across agent tools, so the method can be used beyond Raven.
Testing... testing...
Next, Claude created two fictional organisations, each with a hidden "answer" key:
Tidewell Community Energy, a small co-operative doing rooftop solar, home energy advice, and a schools club. It contains eleven planted problems. These include: a board that talks only about cash flow, forward planning that depends on the chair reading a newsletter "when I get a chance", a volunteer’s near miss on a ladder that never reached the install lead, and a waiting list that more than doubled in eight months. The schools club is healthy on purpose, as a control.
Calder Valley Food Network is a food bank charity with three hubs and a surplus food operation. It contains seven planted problems. One hub is modelled at the next level down, since it runs parcels, a café, and an advice desk in one building. Its coordination, safeguarding route, and trustee visits are healthy, as controls.
Each case has "interview notes" (which are the only material the agent sees), along with an answer key, and a model I built by the agents, as a practitioner would.
Test 1
The first test was of the checker itself. It found 11 of 11 problems in Tidewell and 7 of 7 in Calder, with no false alarms on the healthy controls. It also made two additional findings ("purpose drift") which came out as prompts for human judgement. After all, it's not up to an AI to decide whether an organisation has drifted from its purpose. That's for humans to decide.
Test 2
The second test exposed three gaps in the checker that had been hidden in the first test: no check for a missing "resource bargain", too much noise at lower levels of recursion, and one problem that was reported twice.
Note: A "resource bargain" is the deal between a part of an organisation and the centre – i.e. what that part receives, and what it is accountable for in return. "Recursion" means looking at the same organisation at different levels of detail. A food bank network might be one system at one level; one of its hubs might be a system at the next level down.
...And what did we find?
The agents were better at modelling than expected
Across the runs that produced a valid model, the above-mentioned checker found between 8 and 11 of the 11 Tidewell problems in the agents’ models. The average was 9.5 across eleven models. For Calder, it found between 4 and 6 of the 7 problems, with an average of 5.0 across five models.
Impressively, Raven found and used the skill by itself, with no hints from Claude or me. The specialist agent’s best run found 11 of 11 Tidewell problems.
So what did it miss? In the first round, every agent missed the same problem: nobody checks the quality of the home energy advice visits. When I looked at their models, they had all spotted it. They recorded an audit mechanism described as “No direct audit of Warm Homes visits”. The checker saw an "audit" covering the advice service and moved on. 🤔
The agents understood the organisation. But the schema, the fixed structure the model had to follow, gave them no way to say “this is missing”. They recorded the gap as a thing instead of a thing that was missing.
When Claude added a way to mark a mechanism as absent (via a check that catches descriptions beginning "No…" or "Nobody…") the agents used it properly in the next round.
What did I learn from this? If you build tools for this kind of work, give people, and agents, a clear way to record a gap. Otherwise they will record it as a thing.
Agents make similar mistakes to people
With the Calder example, the agents:
Chose a different level of recursion. They treated parcels, the café, and advice as the network’s operations, rather than the three hubs. That might be defensible, but it hides the fact that the hub manager runs three services, a building, and a fundraising effort on Sunday nights.
Swapped System 2 and System 3*. They put trustee hub visits under coordination and the shared stock system under audit. Humans do this too. 😅
Took "the central team covers rent and staff" as a working resource bargain. Meanwhile, the operations director was saying, "I don’t know what each hub costs to run."
Again, none of these are AI-specific failures, they're just the judgement calls that can make VSM hard.
A good model doesn't guarantee a good diagnosis
A separate reviewer (i.e. another agent which was given the answer keys) marked each written diagnosis. It looked for problems found, false alarms, claims the notes did not support, and whether the method had been followed. Claude spot-checked its marking against the files.
The clearest result came from a run where the specialist agent couldn't read the skill, a setup problem that was fixed later. Its model helped the checker find 9 of the 11 planted problems. Its written report found only 5, and made 12 claims the interview notes did not support. This included phrases in quotation marks, as if from the interviews: "Sue doing too much" and "Poor resource sharing". These were hallucinations: nobody said either of those things.
With the skill in place, the same agent quoted the notes and made three unsupported claims. So the skill’s biggest contribution was not finding more problems. It was keeping the agent honest.
My checker put words in their mouths
Some of the checker’s explanations said things it could not know – e.g. it might say: "Other operations have some direct check on what actually happens." However, the checker can see only the model, not the organisation itself. But the agents copied those sentences into their reports as if they were facts.
I got Claude to rewrite every explanation so it described only what the model records. For example: "The model records audit reaching some operations but none reaching this one. If so, …"
Also, I got Claude to change the skill so that, for every finding, the agent had to give:
Evidence for it
Evidence against it
A Verdict: "supported", "problem with the model", or "needs checking"
Before making that change, the agents labelled the checker’s findings separately but accepted them without testing them. They did this even when their own quotes contradicted the findings. Infuriating. Nearly all the false alarms about the healthy parts of the organisations came from that. 🙄
Did the change help? Not much, to be honest. But the way it failed was instructive.
Only three of the seven reports in the next round used the evidence test at all. In those three, "evidence against" said "none found" 28 times out of 33. False alarms about the healthy parts still got through, marked "supported". And, sometimes, a contradicting quote appeared a few lines further down in the same document. FFS.
Across the 12 marked reports, false alarms did fall slightly, from 1.6 to 1.3 per report. Unsupported claims also fell from 5.8 to 4.6. However, in a sample of this size, those differences are small enough to be chance.
This was partly Raven's fault as it copies an agent’s identity file into its workspace the first time it runs and then... never updates it. So my specialist agent never saw the new instruction. When I got Claude to fix that and run it again, the agent used the evidence test in both reports it wrote. One still said "none found" for 20 of its 22 findings, but the other was the first to spot that one of the checker’s findings was a problem with its own model. It spotted it, then <drumroll> kept the finding anyway.
The bigger lesson remains: tell a model to test each finding against the evidence and you get the form without the substance: "Evidence against: none found". Just like how people tick a box on a risk register.
If we want real scrutiny here, it has to be built into the structure. So one option is a second agent whose only job is to argue against each finding.
What I learned about Raven
Raven is brand new and pre-alpha, and it shows in places. There are some parts which are only documented in Chinese, so I was glad that Claude is multilingual!
Overall? I was impressed.
The good
raven agents new created a working agent and checked its own output, including the manifest, plugin, Python code, and discovery process, before I had configured a model. Setup can run non-interactively, and rrror messages are clear. The "traces" record which model served every call, which turned out to matter quite a lot as it turned out.
The bad
Delegated work landed in the wrong place. In the first round, the sub-agent wrote its files to a new folder next to the one it was given. The host could not find them, decided the sub-agent had failed, and did the whole job itself.
Upstream failures can look like success. Several runs died when a free model returned an error mid-reply. The library Raven uses turns that into a normal stop, so Raven finished with a success code and no files.
Delegation and one-shot runs do not mix well. A sub-agent’s result arrives as a new turn, and a new turn cannot start until the current one ends. In one run, the sub-agent finished in seven minutes. The host then spent the next fifty minutes running
sleep 180in a loop, waiting for a result it could not receive until it stopped waiting.
The ugly:
The free router needs a doubled prefix.
openrouter/freefails;openrouter/openrouter/freeworks. The library underneath strips one prefix.Skills cannot be symlinked into Raven, but agents can. Skills have to be copied, including into each sub-agent’s own folder, or the sub-agent cannot see them.
Adding your own agent hides the built-in ones. A home
agentsfolder replaced the shipped research, code, and on-call agents on the roster.Agent identity files are copied once. Edit an agent’s
soul.mdand the running agent will not see the change. You have to delete its copy first.Some things escape its home folder. The skill cache is hard-coded to
~/.raven, and every working folder gets a.raven/shadow.gitcheckpoint repository.
Free models are cheap and somewhat chaotic
OpenRouter’s free router picks a model per call, not per task. As a single diagnosis typically passed through six to eight different models, from a 550-billion-parameter model down to small “flash” models, things weren't very repeatable. For example, on the very first test call, the router sent "say hi in three words" to a content-safety classifier. 🤦♂️
I got Claude to try pinning single free models to make the runs repeatable. This actually made things worse! One was rate-limited across the free pool and the other kept failing partway through long replies.
The router is chaotic, but it routes around failures, which is good. Across the experiment, 16 of 22 runs on the router produced a valid model. Pinned free models managed 0 of 4. For a real comparison of methods, I would pay for a single pinned model. The free tier was good enough to learn a lot, but not good enough to measure precisely.
Caveats
This was a time-limited test, so longer and deeper experimentation will produce different (and perhaps better) results:
The test cases are not independent. The same AI assistant that helped me build the checker wrote both fictional organisations and their answer keys. The planted problems fit the checks neatly for that reason, whereas a real organisation would be messier.
The marking was done by another AI agent. I spot-checked some of its claims against the files and they held up. However, one marker is still one marker.
The samples are tiny. 26 runs, two cases, and a different mix of free models every time. Differences of one false alarm or a couple of unsupported claims per diagnosis are noise.
I didn't test Raven’s most distinctive features: the Curator, orchestration of several agents together, or its memory. Our agent ended up hiding the built-in ones, and I never fixed that. Something to come back to.
Yeah, so what?
You might have got this far and wonder why on earth anyone would care about this. However, these are my conclusions so far:
An agent can make a credible first pass at modelling an organisation from interview notes, if the model is visible and challengeable, and the method insists on evidence.
The checks should be boring and transparent as their value lies in the questions they prompt, not the verdicts.
The hard judgements are the ones VSM practitioners already find hard: which level of recursion to look at, coordination versus audit, and whether a mechanism really works. The agents struggle in the same places we do.
Honesty needs designing in. Without the skill, an agent invented quotes. With a checker that asserted things, agents repeated them as fact. Every place a tool speaks confidently is a place an agent will copy it.
Would I use this with a real client yet? Not unsupervised. 😅
As a way of turning a stack of interview notes into a first model that a practitioner then argues with, I think it it's pretty close, actually.
Coming next
Pinning a single, paid model and rerun the tests, to separate the method from model quality.
Adding a "challenger" agent whose only job is to argue against each finding, and see whether that does what the written instruction did not.
Running the same skill in a different agent tool, to see how much Raven itself adds.
Trying Raven’s experimental Curator: can a few rounds of practitioner feedback improve the modelling?
Trying it with a real organisation, with consent.
More soon!
For the sake of transparency, Claude created the first draft of this note. Image: Are.na