All jobs

Founding Engineer, Agent Systems

Location
London
Employment Type
Full-time
Location Type
On-site
Level
Mid-Senior Level

About the company

A seed-stage company is building agent-native risk infrastructure: risk management and trust building, delivered by AI agents, for a world increasingly run by them. Their agents sit on top of proprietary data and reassess continuously rather than at fixed checkpoints - so customers spend their time deciding and acting on what matters, not assembling evidence to get there.

Seven-figure revenue within months of launch, on multi-year contracts with leading enterprises in financial services, regulated technology, and healthcare. Founders from Palantir, Oxford, Stanford, and ETH. Backed by leading UK and US institutional investors and angels from Meta, Isomorphic Labs, Palantir, and SpaceX.

The role

You own the agent platform: the orchestration, evals, and reliability work that turns model calls into product features customers trust. The bar isn't that the demo works - it's that a domain expert reading the agent's output considers it at the level of a peer.

This isn't a research role at its core: the team consumes frontier APIs and makes them production-grade. They push them hard - hard enough to have recently found and reported a bug in the Anthropic API that took their engineers weeks to reproduce. At that level, the line between using models and studying them gets thin, so if research-flavoured work pulls at you, there's room to follow it.

What you'll do

  • Agent scaffolding: tool use, context management, sandboxing, prompt-injection defence

  • Evals for fuzzy, high-stakes outputs: assessments, policy interpretation, control mapping

  • Reliability infrastructure: retries, fallbacks, circuit breakers, prompt versioning

  • Set the internal standard for what "good enough to ship" means for AI features

What you bring

  • Backend engineering in TypeScript (or comparable), with 1-2+ years shipping production LLM features

  • Experience with agent frameworks, tool calling, and multi-step orchestration

  • Production evals: dataset curation, LLM-as-judge failure modes, regression testing under model swaps

  • Strong systems thinking: async, queues, idempotency

  • Comfort being the named owner of AI quality, including saying no when needed

Nice to have: Anthropic, OpenAI, or open-weight APIs in production at scale · prompt-injection or agent-security work · background in compliance, audit, or any domain where correctness is fuzzy and stakes are high

Working here

  • King's Cross, London (Gridiron building) - in-person by default, flexibility for days that need it

  • Top decile London market compensation + meaningful EMI-eligible equity

  • Daily team lunch, specialty coffee, roof terrace, on-site showers, serious AI tooling and API budgets

  • Three-stage interview: behavioural phone screen, technical phone screen, paid on-site work trial - under two weeks from first conversation

Stack: TypeScript · Node.js · React · Tailwind · Express · Azure (Container Apps, Service Bus, Front Door, Entra ID) · Postgres · Terraform · GitHub Actions · Docker · Anthropic-first AI · Claude Code throughout

Apply now
Founding Engineer, Agent Systems
Apply now