Senior AI Engineer, Agents
Why Telepatia
Telepatia AI is a healthtech company on a mission to transform healthcare in Latin America and beyond. Born at Stanford and built by doctors for doctors, our platform introduces AI Healthcare Employees that work invisibly inside clinical workflows - transcribing consultations in real time, generating structured medical records, offering evidence-based decision support, and integrating seamlessly with any current health system - all without disrupting the physician's flow.
You'll take clinical AI agents that already run in production and push them to state of the art - ontologies that make decision support deterministic, small models trained to specialize the stack, routers that match task difficulty to compute, and evaluation that finds what standard frameworks miss. Small team, hard problems, and systems that doctors across Latin America use every day.
What you'll do
Design clinical ontologies and structured knowledge representations that agents consume - so decision support becomes deterministic and auditable, instead of "the model usually gets it right"
Fine-tune and distill small LLMs to specialize and accelerate the agent stack, from data curation through deployment
Decide when a problem is solved by better context, when it needs training, and when it belongs in deterministic code - and justify the call on quality, latency, and cost. Context engineering here is measured engineering: retrieval, compression, memory, structure. Not prompt tweaking
Build router models and task decomposition. Make the easy 80% cheap and fast; make the hard 20% correct
Push evaluation past the standard playbook - adversarial and counterfactual evals, calibrated LLM-as-judge, failure taxonomies, uncertainty estimation. Our framework is already good; your job is to find what it misses
Own these systems in production - monitoring, cost, failure modes - and set the bar for how agents are built and evaluated here
What you bring
Senior, hands-on depth in this exact work: models you trained end to end where the result mattered, evaluation systems you designed that caught what standard metrics didn't, agents you shipped and kept running
Clear judgment on context versus training, and when each is the right answer
Fluent in context engineering - RAG, vector stores, knowledge graphs, grounding agents against large-scale data
The ability to look at a clinical rule and decide what belongs in code, what belongs in an ontology, and what belongs in the model - and defend the split
Strong Python. You read research and can tell which results survive contact with production
AI-native: you code with AI tools daily and get real leverage from them
Fluent in English
Bonus
Healthcare or clinical domain knowledge
Experience in fast-paced environments