Terry Li Selected essays
AI controls architecture
I write for practitioners making AI useful and governable inside regulated institutions. These twenty essays examine controls for systems that reason, use tools, and act, then follow those controls into operating practice.
- AI Controls Architecture
Risk teams know risk. The open problem is designing controls for systems that are non-deterministic, probabilistic, and attackable in natural language.
- The Risk Without an Engineering Solution
Every other agentic AI risk has an engineering answer. Prompt injection doesn't. That changes everything about how you design controls.
- Why Agents Break Governance
Four interactions between agentic properties create risks that manual governance cannot address. The category boundary is not AI versus traditional — it is systems that act versus systems that advise.
- Govern the Workflow, Not the Model
Agent governance cannot stop at model behavior. Once AI systems use tools, the governed object is the whole workflow.
- The Agent Is Not the Control Point
Finance agents are evidence custody systems before they are model systems.
- The Calibration Gap: Why AI Assurance Needs Experimental Rigour at Consulting Speed
AI governance keeps choosing between rigorous-but-slow validation and fast-but-unfalsifiable frameworks. The escape is the experiment itself: for any measurable claim, the measurement is the faster path.
- The Judgement Bucket: How to Be Rigorous About the Claims You Can't Measure
The sequel to the calibration gap: some governance claims have no number, and that does not excuse them from evidence. Pre-register the rule, invite the refutation, name the residual, report honestly.
- Your Check Passed on the Proxy
A green result is a true statement about a stand-in. Whether that stand-in resembles the thing you care about is a separate question, and the check cannot answer it.
- Your Reviewer Model Is Not Independent
When agents generate faster than anyone can read, the standard answer is a second model reviewing the first. Aerospace decomposed what makes a check independent decades ago, and a reviewer model fails the hardest of the three tests.
- Retrieval Becomes a Control
Cerebras published an unusually honest account of their internal knowledge base. The retrieval engineering is solved. The assurance layer gets one sentence. Every enterprise copying this pattern is importing that asymmetry.
- Governing Agents the Way Cells Govern Themselves
Six cell biology mechanisms that reveal what the networking 'control plane' metaphor misses about governing AI agents.
- A Skill Is Not a Prompt
The useful unit in agent systems is not a better instruction. It is a tested capability package: judgment, code, checks, routing, and boundaries.
- The Program Is the Plan
For many agent workflows, the right abstraction is not another tool call. It is a bounded program over small primitives.
- The Agent Is the Trace
Long-running agents are not defined by the model call. They are defined by the state, rules, tools, failures, and corrections that survive it.
- The missing layer between model risk and application security
Model risk reviews the model. Application security reviews the application. Neither sits behind the agent at execution time, watching the verbs as they go out.
- After the Harness
Once model companies supply the generic agent harness, the valuable work moves into workflow design, human intervention, domain data, and the definition of good work.
- The SOP Is the Product
Enterprise AI stops being a chatbot when the operating procedure becomes the thing the system can execute, inspect, and improve.
- The Model Is Not the Unit of Return
Model revenue is not customer return. The economic and risk unit is the harness that turns model output into accountable work.
- When Code Gets Cheap, Coordination Gets Expensive
Coding agents move the bottleneck from implementation to shared intent.
- After Automation, Judgment Becomes Infrastructure
When execution gets cheap, the scarce work moves to framing, review, and the systems that preserve judgment.