Skip to content

AI workflow · 15 posts

The model is 30% of the win.

I ship production features across a dozen repos with AI agents every day. The model matters far less than the harness around it — the verification gates, the deterministic hooks, the cheap always-on watchers, and the context architecture. This is the whole setup, warts and all, plus every way it has burned me.

The models and the money

Which model to point at the problem, what it costs, and why the answer keeps changing every quarter.

How AI agents actually fail

Named failure modes from real sessions — unverified "done" claims, confident agreement, and building the wrong thing thoroughly.

Building the harness

The model is maybe 30% of the win. This is the other 70% — context architecture, deterministic guardrails, cheap watchers, and multi-agent orchestration.

Let us make some quick suggestions?
Please provide your full name.
Please provide your phone number.
Please provide a valid phone number.
Please provide your email address.
Please provide a valid email address.
Please provide your brand name or website.
Please provide your brand name or website.