Agent hardening for production
You can ship an agent in a sprint. Building the controls to run one takes longer. Gartner expects over 40 percent of agentic AI projects to be canceled by the end of 2027, and names inadequate risk controls among the causes. The ones that survive will be the ones someone hardened.
I’m Saša Tomić, and I harden agents for production: red teaming, eval gates in CI, identity and permissions, and the switches that stop a misbehaving agent before your customers notice. At DFINITY I put production AI in front of real users and kept a 1,400-node network reliable, so I have felt what an uncontrolled component costs at 3am.
Who this is for
Maybe you already run agents: a support bot with tool access, an internal copilot reading company data. Your team built the demos, and the demo worked. Now procurement or a customer questionnaire asks who the agent authenticates as, what it can touch, and how you stop it. If your team cannot answer in one sentence, this is for you.
What you get
The offer is three builds. Every one ends with controls running in your systems.
1. The red team, before a stranger runs it for you
- I attack your live or staging agents: prompt injection, jailbreaks, data exfiltration, runaway tool use, permission overreach.
- You get a written attack report: how far the attack got and what it would have cost you.
- I implement the fixes, verify them by re-test, and close the findings, so nothing ends up decorating a slide deck.
2. Evals in CI, so a regression cannot merge
- I build golden test sets from your real tasks and failure cases.
- Your reviewers calibrate the judge scorers on the first batches, so “good” means what your product needs.
- The gate is merge-blocking: it re-runs the attack suite and the task suite on every prompt, model, or tool change before that change ships.
- You also get a go/no-go scorecard per agent, so shipping decisions have numbers behind them.
3. Identity, permissions, and the brakes
- I inventory every agent and the credentials it carries.
- Least-privilege permissions per agent, spend caps per task, and an audit trail of agent actions.
- You get kill switches and human-approval gates on the actions that matter: payments, deletions, outbound messages.
- A one-page runbook: who pulls the switch, what happens to in-flight work, and how the agent comes back.
Why me
On Caffeine.ai I built the AI review loop that kept 5 to 15 pull requests merged per engineer per day, so I have run agents against real code at production volume. The rest is on the about page.
What hardening looks like
Two weeks of red teaming and inventory, then I land the eval gate and the controls in your CI and your platform. Remote and async, NDA first. I work in your repository with least-privilege access, and every artifact stays with you. Fixed scope, fixed price, quoted on a half-hour call.
Questions I hear
“Our agents passed a penetration test.” A network pentest probes your perimeter. An agent attack starts inside it: a crafted document tells your bot to forward the customer list somewhere new, and a port scanner cannot see it.
“We already have observability.” Your dashboards show what the agent did. An eval judges whether it was right, and the gate in CI blocks the regression before it ships.
“Isn’t this the model vendor’s job?” The vendor owns the model. You own what it can reach: your data and your customers. You enforce that boundary with controls in your own systems.
“We only use one agent framework.” Good. I wrap the hardening around whichever framework you run, and the eval gate keeps working when you switch.
Next step
My LinkedIn inbox is open: name the agents you run and what they can touch. Within two days you have a written read of your worst exposure, even if the answer is that you do not need me.
Sources
- Gartner, press release, June 2025: over 40 percent of agentic AI projects forecast to be canceled by end of 2027, citing escalating costs, unclear business value, or inadequate risk controls.
- Cloud Security Alliance, 2026: 53 percent of organizations have had AI agents exceed their intended permissions.
- IBM, Cost of a Data Breach Report 2025: 97 percent of AI-related breaches involved systems lacking proper AI access controls; high levels of shadow AI added an average of $670,000 to the cost of a breach.