PharmaTools.AI — Nick Lamb · AI Research & Engineering

Evaluation, verification and oversight for autonomous AI systems.

AI agents are starting to do real work — and mostly, we take their word for it. I design experiments that reveal how these systems actually behave, and build open-source tools that check their answers against evidence.

agents & multi-agent systems·evaluation·evidence & verification·privacy·healthcare AI

Nick Lamb
Nick Lamb, PhD, CMPPIndependent AI researcher & engineer · 20+ years in regulated medical communications

01Flagship research · Observer Zero

What happens when AI agents become scientists — inside a world whose ground truth we control?

Observer Zero is a persistent artificial society for studying how autonomous LLM agents gather evidence, interpret it, revise their beliefs and reach collective conclusions. Its agents inhabit Meridian — a world with fictional physics — where they run experiments, exchange letters and hold explicit, probabilistic beliefs about laws they cannot see. Because the simulator holds perfect ground truth, every belief can be scored against reality: the design separates what a society had the evidence to know from what it actually came to believe.

When the laws of their world secretly changed, the agents gathered evidence enough for a blind statistical detector to find the change — and almost none of them ever believed it.

235seeded simulated universes across two studies
40 / 40eight-agent runs where a detector, fed only the agents’ own measurements, found the hidden change
1 / 276agents concluded that a physical law of their world had changed
6,880 / 0agent-days of opportunity in homogeneous grounded societies — and zero voluntary letters
OBSERVER ZERO · OUTSIDE THE WORLD observations letters MERIDIAN fictional physics · hidden laws g = 14.20 → ? Ada Maya Theo Samuel Elena Leah Tom Jamie EVERY OBSERVATION, LETTER AND BELIEF — LOGGED AND SCORED

Eight agents measure a world they can never see into directly. Observer Zero sits outside it, holding the ground truth — so what each agent could know can be separated from what it came to believe.

STUDY 01Complete

When the Laws Changed

Can autonomous agents recognise that a physical law of their world has changed?

Across 150 universes, 90–100% of agents transiently noticed something was wrong — yet 0 of 40 gravity-shift worlds reached a strict law-change conclusion. The anomalies were real; the interpretation reached for faulty instruments instead.

STUDY 02Complete · under review at JASSS

Does Society Help?

Can an artificial scientific society overcome the individual failure through social interaction?

Pre-registered, 85 universes, societies of two and eight. Left alone, grounded societies never spoke; seeding one communicative agent spread unsupported claims, not scrutiny. The combined two-study paper is under review at JASSS, with the model published in the CoMSES library.

STUDY 03Next · in design

The Simulation Question

What would it take for an AI scientist to discover that its world is simulated?

The next study, now being designed, targets world-model revision: when does an intelligent system stop assimilating anomalous evidence into its existing explanation and conclude that the explanatory model itself is inadequate?

From research questions to working systems.

Strong claims about AI behaviour need trustworthy measurement. One question runs through all three of these systems — how do we know an AI is doing what we think it is? — and each answers it in a different domain. Every one ships with its own evaluation.

Evaluation & verification · open source

OpenGATE

Can an AI claim prove where it came from?

An open-source verification framework for evidence-grounded AI. Deterministic grounding and extraction checks — no LLM judge anywhere in the loop — so verdicts are reproducible, auditable, and cheap enough to run on every answer or every commit. Four production systems run on it in CI, where it has caught silent parse failures, a citation-parser inconsistency, and a simplifier dropping an antibiotic dose from a discharge summary.

npm · 1.9K downloads/moPyPIMCP serverGitHub Action
Privacy infrastructure · Anthropic MCP Directory

Redacta

How can agents work with sensitive information without exposing it unnecessarily?

A privacy boundary for AI and agent workflows. Redacta pseudonymises patient identifiers — names, NHS numbers, dates — before text reaches a model, and its v2 MCP server keeps the reversal map outside the model’s context entirely. Benchmarked under adversarial attack by its own published evaluation harness, and deployable from a free iPhone app to a self-hosted Kubernetes cluster.

iPhone appMCP serverSelf-hostedAdversarially benchmarked
Production healthcare AI · MHRA Class I device

Patiently AI

Can generative AI simplify medical information without losing meaning?

A deployed, MHRA-registered Class I medical device that turns clinical letters, results and discharge summaries into plain language — in 18 languages, on web, iOS and Android. Validated in a three-phase study — computational readability analysis, expert review by 15 healthcare professionals who rated outputs 4.49/5 for accuracy, and a 54-patient survey. Five national awards, in real-world use.

MHRA Class I18 languagesValidation study5 national awards

Who’s behind this.

I build AI systems — then ask harder questions about whether we should trust them.

Nick Lamb is an independent AI researcher and engineer with an unusual route in: twenty years of regulated medical communications (PhD, CMPP), a field where every claim must trace to evidence. That standard now runs through everything on this page — agents whose beliefs are scored against hidden ground truth, verification that never takes a model’s word for it, and a registered medical device validated before it reached patients.

PharmaTools.AI began as a collection of healthcare AI tools, and healthcare remains its proving ground: few domains punish an ungrounded claim faster. The work now spans empirical AI research, autonomous agents, evaluation and privacy infrastructure — much of it open source.

Interested in the same questions?

How autonomous AI systems reason, use evidence — and fail. If you’re working on related problems, or want this kind of discipline applied to your own AI systems, I’d be glad to hear from you.