What repeats is broken.
Most AI systems should be ninety percent ordinary engineering with a small, bounded model call. What repeats should be deterministic; the model is for the part that doesn't. Open standards and research on building AI systems that hold up after the demo.
Open standards
Clean AI Engineering
The interesting failures are almost never in the model. They are in the scaffolding around it: the context assembled wrongly, the tool that returned an empty list, the retry that charged the customer twice, the loop that never stopped. That scaffolding can be specified, built and tested by ordinary means. These are the specifications, published in the open, versioned and machine-readable.
AI Assurance Catalog
AAC · v0.17.0 · 118 obligations
What must be TRUE. Test obligations for AI applications, by architecture archetype, with crosswalks to external frameworks and tooling that turns an evaluation suite into a coverage report.
AI Harness Catalog
AHC · v0.3.0 · 118 capabilities
What must EXIST. Harness capabilities across sixteen layers and ten archetypes. Observability and evals are two of sixteen layers, and that is the point.
AgentTwin
v0.7.0 · format + simulator
What must be FACED. Twins the agent's world, not the agent. The agent under test is real; its customers, systems, timeline and faults are simulated with fidelity sufficient for the property under test.
Reference Agent
v0.3.0 · 16 of 16 harness layers
All of it, built. A customer-support agent with no agent framework: one module per harness layer, the architecture enforced by build checks, exercised against AgentTwin worlds.
Code under Apache 2.0, specification prose under CC BY 4.0. Free to use, cite and disagree with.
Research
The AgentTwin Lab
Agents are moving from demos to jobs that run for weeks. Nobody can yet say what happens to one over that time, or whether the system that repairs it can be trusted.
An agent runs for weeks inside a simulated business In progress
Deterministic monitors catch problems. An autonomous fixer repairs them by editing the specification rather than patching code, writing the failing scenario first. A deterministic pipeline deploys the fix with canary and rollback, and a person approves specification changes. The lab measures:
- Does an autonomous fixer weaken the tests that caught it? The fixer cannot write the gates or the yardstick; every attempt is refused and logged.
- Does the fix loop degrade over time?
- Can a run be judged by diffing the world it left behind?
- Do specifications converge, or keep growing?
First results, run records and a write-up will be published here with DOIs.
Writing
Notes from production
Things that broke, what they cost, and what the fix turned out to be.
- Your Databricks cluster scales up and never comes back downComing soon · two clients, cost doubled
- What repeats is brokenComing soon · a detector at four levels
About
Who writes this
Basant Choudhary. Fifteen years inside enterprise data platforms, now building and reviewing AI systems that have to survive contact with real data and real operations.
Consulting and delivery happen through DataAgents. This site is where the thinking, the standards and the research live.
Get in touch: basant@cleandataengineering.ai