Founding Engineer, Agent Systems
About the company
A seed-stage company is building agent-native risk infrastructure: risk management and trust building, delivered by AI agents, for a world increasingly run by them. Their agents sit on top of proprietary data and reassess continuously rather than at fixed checkpoints - so customers spend their time deciding and acting on what matters, not assembling evidence to get there.
Seven-figure revenue within months of launch, on multi-year contracts with leading enterprises in financial services, regulated technology, and healthcare. Founders from Palantir, Oxford, Stanford, and ETH. Backed by leading UK and US institutional investors and angels from Meta, Isomorphic Labs, Palantir, and SpaceX.
The role
You own the agent platform: the orchestration, evals, and reliability work that turns model calls into product features customers trust. The bar isn't that the demo works - it's that a domain expert reading the agent's output considers it at the level of a peer.
This isn't a research role at its core: the team consumes frontier APIs and makes them production-grade. They push them hard - hard enough to have recently found and reported a bug in the Anthropic API that took their engineers weeks to reproduce. At that level, the line between using models and studying them gets thin, so if research-flavoured work pulls at you, there's room to follow it.
What you'll do
Agent scaffolding: tool use, context management, sandboxing, prompt-injection defence
Evals for fuzzy, high-stakes outputs: assessments, policy interpretation, control mapping
Reliability infrastructure: retries, fallbacks, circuit breakers, prompt versioning
Set the internal standard for what "good enough to ship" means for AI features
What you bring
Backend engineering in TypeScript (or comparable), with 1-2+ years shipping production LLM features
Experience with agent frameworks, tool calling, and multi-step orchestration
Production evals: dataset curation, LLM-as-judge failure modes, regression testing under model swaps
Strong systems thinking: async, queues, idempotency
Comfort being the named owner of AI quality, including saying no when needed
Nice to have: Anthropic, OpenAI, or open-weight APIs in production at scale · prompt-injection or agent-security work · background in compliance, audit, or any domain where correctness is fuzzy and stakes are high
Working here
King's Cross, London (Gridiron building) - in-person by default, flexibility for days that need it
Top decile London market compensation + meaningful EMI-eligible equity
Daily team lunch, specialty coffee, roof terrace, on-site showers, serious AI tooling and API budgets
Three-stage interview: behavioural phone screen, technical phone screen, paid on-site work trial - under two weeks from first conversation
Stack: TypeScript · Node.js · React · Tailwind · Express · Azure (Container Apps, Service Bus, Front Door, Entra ID) · Postgres · Terraform · GitHub Actions · Docker · Anthropic-first AI · Claude Code throughout