OrgLens: 3. Teaching a tool to be careful, then giving it more than one lens

Notes updated ~6 min

Gif showing dots in a patternThe first two OrgLens notes (1, 2) asked whether AI could turn what people say about an organisation into a useful health check.

The results made me cautiously optimistic: OrgLens could produce a structured description of an organisation and a plain-language report. It did, however, need software checks to verify quotations, catch contradictions, and stop important issues disappearing during editing.

An organisational health check needs to do more than describe who does what. It should help people discuss why the same problems return, where their views differ, and whose experiences are missing.

So this third note covers those additions, as well as exploring whether an AI tool could learn to build its own safeguards from feedback.

For those who haven't read the previous notes, I set the questions and decided what counted as a useful result. Claude Code turned those decisions into instructions for Raven, which wrote code, ran tests, and produced trial reports. I reviewed the results and decided what to change.

Can feedback improve the tool?

Raven’s experimental "Curator" feature changes an agent’s instructions in response to feedback. It can add code that checks the agent’s work before anyone receives it.

I tested it on a guided interview: someone describes their organisation, the agent asks up to two questions at a time, then writes a report. We used two fictional organisations for training, with a third kept back to test whether the changes worked on an unfamiliar case.

I compared a version trained over three rounds with an untrained version using the same starting instructions. The interviewees and assessors were AI models, not people.

The Curator received our handbook and interview questions, but none of the checks developed earlier. It created:

  • Planner to track which topics had been covered and which issues had been raised.
  • Checks to:
    • enforce the two-question limit.
    • ensure each recorded issue received its own point in the report.
    • match quotations to an interviewee’s words.

After the first training round, the trained version never exceeded two questions per message. The untrained version broke that rule in every interview. In the final assessment, the trained version met 25 of 27 standards; the untrained version met 23.

Finding problems barely improved: 24 of 26 planted problems against 23. That difference is too small to support much of a claim. Training, then, made the agent more disciplined. It didn't show better "judgement".

There was also a cost to stricter checking: for example, one check rejected eight report drafts in succession until the agent ran out of time. A safeguard needs a way to stop retrying and explain what remains unresolved, rather than leaving someone waiting indefinitely.

Beyond organisational structure

OrgLens started with the Viable System Model (VSM), which examines how an organisation coordinates work, checks what is happening, looks ahead, and responds to trouble.

I wanted to add two other ways of looking that go beyond VSM:

  1. System Dynamics – feedback loops: circles of cause and effect (e.g. a growing waiting list might increase staff pressure, leaving less time to deal with the waiting list). This also looks for recurring patterns, called "system archetypes", such as a quick fix that makes a longer-term problem worse.
  2. Perspectives – what different people think the organisation is for, where those views conflict, and who is affected but has not been heard. This draws on Soft Systems Methodology (SSM) and Critical Systems Heuristics (CSH), two approaches that question how people define a situation.

Together, these lenses ask:

  • How is the organisation set up?
  • Why do problems keep returning?
  • Do people want the same things?
  • Whose views are missing?

I asked Claude to create answer keys for the fictional organisations before we built the new lenses. We then developed specifications, which Claude passed to Raven-Code, Raven’s coding agent.

What the tests showed

I tested all three lenses (VSM, SSM, CSH) on three fictional organisations over four rounds, using OpenRouter's free AI model router.

Combining the findings took work. The same issue could appear several times, under different lenses. An existing contradiction check stopped working when the report structure changed. In another run, the software accepted an empty template from one lens, yet the report still sounded convincing using results from the other two. Annoying.

By the fourth round, all three lenses produced results for all three organisations. Every quotation was checked against the source material, and the reports contained no systems-thinking jargon.

The perspectives lens found most conflicting views and missing voices. Feedback loops were less consistent: the tool found 24 of 36 expected loops across the four rounds. Wider recurring patterns were weaker still, appearing in only 6 of 20 expected cases.

For now, OrgLens is better at recording a circle of cause-and-effect than reliably recognising a familiar pattern within it.

Whose idea of “better”?

CSH asks who a situation serves, who decides, whose knowledge counts, and who speaks for people affected but not involved. I can be used both as a diagnostic ("is") and goal ("ought").

With the food bank, the model supplied answers to all 12 "ought" questions. Most were its own recommendations, but presented as if they reflected the organisation’s views.

I asked Claude to instruct Raven to require a quotation supporting every statement about what should happen. That reduced the answers from 12 to 6, but didn't solve the problem.

Three matched the answer key. One was partly supported. Two still used descriptive quotations to justify recommendations. “Nobody saw it coming” describes what happened; it does not tell us what the speaker thinks should happen.

The tool missed a purpose gap elsewhere too: the energy co-operative said it served “everyone in the town”, but most work went to schools and community buildings.

Checking that words appear in the notes is not the same as checking that they support the conclusion. OrgLens can show the evidence. People still need to judge what it means.

What this means in practice

These tests involved fictional organisations, simulated interviewees, and AI assessors. The same AI setup helped build the tools and evaluate them. They are useful experiments, not evidence that OrgLens is ready to diagnose real organisations independently.

The additions make it more useful as a starting point for discussion. They can help people examine recurring pressures, compare different accounts, and ask who else needs to be heard.

They do not give the tool authority to decide what an organisation should be for, which view is right, or whether its arrangements are fair.

My next step is to hear from several people in an organisation and test OrgLens with real participants who have given consent. Before adding more approaches, I want to know whether these questions help people have a better conversation.


For the sake of transparency, Claude created the first draft of this note. Image: Are.na.

Never shown publicly, used only for Gravatar