AI Product Engineering
AI features that work past the demo.
For teams adding agents, search or recommendations to a real product, or whose AI prototype is too slow, too costly or too unpredictable to ship.
- Signs you need this
- The demo looked great, but real users get wrong or made-up answers.
- Your LLM bill is growing faster than your usage.
- You can't tell whether a prompt or model change made things better or worse.
- You're not sure whether you need agents, RAG, fine-tuning, or none of them.
- What I do
- Architecture review. Pick the simplest approach that works, including telling you when a SQL query or a few rules would beat an LLM.
- Build. Agents that call your APIs and stop for a person to approve anything risky, RAG and hybrid search on Postgres and pgvector, often with no separate vector database to run, and recommendation engines.
- Evaluate. An eval harness of real test cases, scored automatically, that blocks a deploy when quality drops. Plus dashboards and alerts on answer quality, latency and cost in production.
- Control cost. Send each request to the cheapest model that handles it well, fall back to another provider when one goes down, and set hard spend limits.
- You get
- AI features with a test score behind every release and a cost you can forecast, built on the database and cloud services your team already knows.
- Typical shape
- An architecture review with a written recommendation, or a fixed-scope build.
- Why me
- At Goodword I shipped an LLM agent with human-in-the-loop review, hybrid search on Postgres and pgvector, and a recommendation engine to production, with an eval harness that every deploy had to pass. Years in healthcare and payments, where a wrong number gets audited, shaped how I test systems that don't always give the same answer twice.
Next step
You'll get a straight answer on what to fix first, whether or not we work together.