Senior ML Engineer
About the company
An AI/ML tech startup developing a novel foundation model with a singular vision: to achieve fully automated, unsupervised software delivery in embedded control systems. Backed by venture capital and scaling fast.
The role
We're looking for a Senior ML Engineer to lead the research, development, and production deployment of our foundation model - defining the long-term technical strategy and owning ML performance end-to-end.
CUDA proficiency is essential for this role. You'll be designing and implementing GPU-accelerated components including custom CUDA kernels where off-the-shelf libraries fall short. If you don't have deep, hands-on CUDA experience, this isn't the right fit.
What you'll do
Lead research, development, and production deployment of the foundation model
Define the long-term technical strategy for high-performance ML systems
Optimise the model for peak performance across diverse hardware and ensure scalability for exponential user growth
Serve as technical guardian for the model's quality and SLOs
Provide hands-on solution architecture for the core ML infrastructure
Select, evaluate, and implement state-of-the-art technologies (distributed training, specialised hardware, efficient serving frameworks)
Profile and optimise the end-to-end ML stack: data pipelines, training loops, inference serving, and deployment
Design and implement GPU-accelerated components, including custom CUDA kernels
Work closely with the founders to translate product requirements into concrete optimisation goals and technical roadmaps
Build internal tooling, benchmarks, and evaluation harnesses
What you bring
Technical (essential):
Deep, hands-on CUDA C/C++ proficiency - this is a hard requirement
Expertise in designing, architecting, and implementing large-scale foundation models
In-depth knowledge of recent architectures (e.g. MoE, state-space models)
Understanding of data curation and quality control for massive training datasets
Significant hands-on experience optimising and debugging deep learning models
Practical experience with distributed or large-scale training and inference
Deep understanding of at least one major deep learning framework (ideally PyTorch)
Proficiency in Python
Experience building and operating ML systems on cloud platforms (AWS, Azure, GCP)
Containerisation, orchestration, experiment tracking, monitoring, evaluation pipelines
Mindset:
Passion and determination
Able to grind through complicated and ambiguous problems
Delivery-oriented - respects timelines and commitments
Openness to disagreement
Location
This role is open to remote candidates based in Brazil. For candidates interested in relocating to the UK, we are open to supporting visa sponsorship.
About the interview
No live or take-home coding tasks. Maximum 2 interviews: 1hr cultural (online) with the CEO and 2hr technical (face-to-face) with the CPO. Ask anything - full honesty is welcomed.
Why join
Significant impact from day one on a never-before-seen foundation model
Unusual transparency - no management jargon, simple truth
Share options + competitive salary