All jobs

Founding Machine Learning Engineer

Location
Worldwide (remote)
Employment Type
Full-time
Level
Mid-Senior Level

Own the video data pipeline that frontier robotics labs train on, from raw capture to validated episodes.

Hub is one of the fastest-growing data companies, providing real-world training data to the largest AI and robotics companies. Based in San Francisco and backed by Y Combinator and top VCs, our mission is to advance embodied AGI through real-world data and research.

The Role

As a founding ML engineer you own Hub's video data pipeline end to end, from raw capture in the field to a dataset a frontier robotics lab trains on. It is the layer that decides what we are allowed to ship.

What You'll Own

  • Petabytes of video from multi-camera rigs we build ourselves: RGB-D, RGB and IMU, metric depth, global and rolling shutter, hardware sync. You turn raw bundles into validated episodes.

  • Egocentric video with narration, across languages and environments: speech recognition, translation, chapter segmentation, QA against each customer's taxonomy.

  • Quality control as an ML problem. You benchmark vision-language models, decide which ones we trust to judge our data, fine-tune and distill our own, and own the eval sets and thresholds behind that call.

  • Scalable QC and QA pipelines that trim, cut and quarantine, so no clip ever ships out of spec.

  • Annotation at scale against demanding customer taxonomies, plus hand tracking and fine-grained manipulation. Every human verdict becomes a training label.

  • The next modalities: tactile and teleoperation sit in the same problem space.

  • Live production for several of the top 5 AI companies, at high volume, with hard deadlines and specs.

The Profile

  • A top engineering school, 3+ years of applied ML in computer vision or robotics. Less experience is fine for outliers: the bar is what you've built.

  • You've run VLMs as judges in production: panel agreement, calibration against human labels, fine-tuning and distillation.

  • You've trained and deployed CV or robotics models in production: detection and tracking, pose and hand estimation, depth, visual-inertial odometry, action recognition.

  • You know the video stack deeply: codecs and frame timing, multi-view geometry, intrinsics and extrinsics, distortion, temporal alignment across sensors.

  • Exceptional individual achievement: elite rankings at competitions or concours, hackathons won, projects at real scale.

  • An active GitHub and Hugging Face: recent contributions, open weights and datasets, reproduced results.

  • You follow the literature and can tell what's worth implementing from what's noise.

  • Agentic engineering as a craft: a custom harness, and a loop for your agents to verify their own work through tests, training runs and evals.

Nice to Have

  • You can come to our Paris office once it opens.

  • Egocentric vision, IMUs, MCAP, ROS.

  • World models, video generation, VLAs or robot foundation models.

Stack

  • PyTorch, CUDA, multi-GPU training and distributed inference. Quantisation, batching and throughput matter as much as accuracy.

  • VLMs as judges: vLLM serving open-weight models on H100s, hosted models behind one provider interface, and our own fine-tuned and distilled judges.

  • Multi-stream HEVC video plus high-rate IMU per recording, hardware-synchronised and calibrated, delivered as MCAP.

  • Postgres for state, S3 for bytes, Kubernetes for compute, GPU inference at petabyte scale. Cost per hour processed is an engineering target.

  • Every threshold is a named constant tied to the customer requirement it comes from. Every quarantine carries a code, evidence and an owner.

What We Offer

  • $90,000 to $120,000 yearly salary.

  • Stock options between 0.25% and 0.5%, granted at signature.

  • Your own GPU budget.

  • Being able to come to our Paris office once it opens is a big plus.

  • Direct work with the founders and with the biggest AI labs.

How We Hire

  1. A short application read by the team. We open your GitHub and Hugging Face first.

  2. A technical conversation with one of our ML engineers.

  3. A build task on real data from our pipeline. We watch how you break the problem down, what you measure and how you verify your own work.

  4. A final conversation with a founder and an ML engineer.

  5. An answer within 48 hours.

How to Apply

Send a short application in your own words:

  • Your CV

  • The achievement you're most proud of

  • Your GitHub and Hugging Face links

  • The story behind it

Apply now
Founding Machine Learning Engineer
Apply now