Guides
Practical guides for teams building with AI.
Written from what we've built, and what broke along the way. Each one starts with a short version, so you can get the point in thirty seconds.
Coding agents
For engineering leaders getting real work out of AI coding agents.
How to verify agent-written code before production
When one coding agent writes the code, the tests, and the summary, all three can share the same wrong assumption and still pass. Give reviewers evidence that does not come from that agent: a test shown to fail when the behavior breaks, a separate reviewer session, and an isolated environment. Start with one release-critical behavior, using the free kit.
Why coding agents break legacy code, and how to fix it
Coding agents break legacy code because it has no tests to tell them a change is wrong. Before an agent refactors, map the dependencies, pin today's behavior with characterization tests, and change one seam at a time. Then the agent speeds up the mechanical work instead of guessing.
Write the spec first: coding agents need more than a prompt
A prompt tells a coding agent what to build now, not what must stay true. For production work, write a short spec first: scope, interface, limits, failure cases, and how you will check the result. Then build in small slices and review each one against the spec.
AI in your product
For product teams shipping AI features that hold up with real users.
Why your AI agent works in the demo and breaks in production
A demo proves the model can do a task once, on clean inputs, while someone watches. Production needs the system around it: saved state, checked outputs, focused context, logs, runs that survive restarts, and customer data kept apart. Design for these from day one; they are hard to add later.
How to check an AI agent's work before your customers do
An AI agent can sound right and still do the wrong thing, so 'it seems to work' is not a test. Check it in layers: hard checks on what it did, rubric scores for what it wrote, business-rule checks, and human review of samples and edge cases. Rerun the same cases after every change.
Why your RAG pipeline keeps disappointing
RAG works well for looking things up in stable documents. It disappoints when the answer depends on live state (this account's stage) or history (last week's call), which document search cannot reliably find. Use three layers: document search, direct lookups for current state, and a short memory summary.
Automating operations
For business leaders deciding what to hand to AI first.
Your team uses ChatGPT. Your workflows haven't changed.
Staff using ChatGPT makes individuals a little faster, but the business still runs the same way. Real change comes from rebuilding one recurring workflow so AI does the routine steps and a person approves what matters. Start narrow, for example matching supplier invoices to purchase orders, and prove it works.
What the first two weeks of AI automation should produce
The first two weeks should get one real workflow running on your own examples, with a person approving the output and an honest list of what worked and what failed. Agree on scope, what done means, and when you would stop before any work starts.
Tell us what's slowing your team down.
You'll talk to Nadeem, our founder, and get a straight answer on what AI can and can't do for your team.
If the first two-week phase doesn't satisfy you, you don't pay for it.