Founding Machine Learning Engineer
Own the video data pipeline that frontier robotics labs train on, from raw capture to validated episodes.
Hub is one of the fastest-growing data companies, providing real-world training data to the largest AI and robotics companies. Based in San Francisco and backed by Y Combinator and top VCs, our mission is to advance embodied AGI through real-world data and research.
The Role
As a founding ML engineer you own Hub's video data pipeline end to end, from raw capture in the field to a dataset a frontier robotics lab trains on. It is the layer that decides what we are allowed to ship.
What You'll Own
Petabytes of video from multi-camera rigs we build ourselves: RGB-D, RGB and IMU, metric depth, global and rolling shutter, hardware sync. You turn raw bundles into validated episodes.
Egocentric video with narration, across languages and environments: speech recognition, translation, chapter segmentation, QA against each customer's taxonomy.
Quality control as an ML problem. You benchmark vision-language models, decide which ones we trust to judge our data, fine-tune and distill our own, and own the eval sets and thresholds behind that call.
Scalable QC and QA pipelines that trim, cut and quarantine, so no clip ever ships out of spec.
Annotation at scale against demanding customer taxonomies, plus hand tracking and fine-grained manipulation. Every human verdict becomes a training label.
The next modalities: tactile and teleoperation sit in the same problem space.
Live production for several of the top 5 AI companies, at high volume, with hard deadlines and specs.
The Profile
A top engineering school, 3+ years of applied ML in computer vision or robotics. Less experience is fine for outliers: the bar is what you've built.
You've run VLMs as judges in production: panel agreement, calibration against human labels, fine-tuning and distillation.
You've trained and deployed CV or robotics models in production: detection and tracking, pose and hand estimation, depth, visual-inertial odometry, action recognition.
You know the video stack deeply: codecs and frame timing, multi-view geometry, intrinsics and extrinsics, distortion, temporal alignment across sensors.
Exceptional individual achievement: elite rankings at competitions or concours, hackathons won, projects at real scale.
An active GitHub and Hugging Face: recent contributions, open weights and datasets, reproduced results.
You follow the literature and can tell what's worth implementing from what's noise.
Agentic engineering as a craft: a custom harness, and a loop for your agents to verify their own work through tests, training runs and evals.
Nice to Have
You can come to our Paris office once it opens.
Egocentric vision, IMUs, MCAP, ROS.
World models, video generation, VLAs or robot foundation models.
Stack
PyTorch, CUDA, multi-GPU training and distributed inference. Quantisation, batching and throughput matter as much as accuracy.
VLMs as judges: vLLM serving open-weight models on H100s, hosted models behind one provider interface, and our own fine-tuned and distilled judges.
Multi-stream HEVC video plus high-rate IMU per recording, hardware-synchronised and calibrated, delivered as MCAP.
Postgres for state, S3 for bytes, Kubernetes for compute, GPU inference at petabyte scale. Cost per hour processed is an engineering target.
Every threshold is a named constant tied to the customer requirement it comes from. Every quarantine carries a code, evidence and an owner.
What We Offer
$90,000 to $120,000 yearly salary.
Stock options between 0.25% and 0.5%, granted at signature.
Your own GPU budget.
Being able to come to our Paris office once it opens is a big plus.
Direct work with the founders and with the biggest AI labs.
How We Hire
A short application read by the team. We open your GitHub and Hugging Face first.
A technical conversation with one of our ML engineers.
A build task on real data from our pipeline. We watch how you break the problem down, what you measure and how you verify your own work.
A final conversation with a founder and an ML engineer.
An answer within 48 hours.
How to Apply
Send a short application in your own words:
Your CV
The achievement you're most proud of
Your GitHub and Hugging Face links
The story behind it