datadoo doesn't ship trained models. We generate the data, and whoever trains on it owns the model. Internally we train evaluation models to measure how well a dataset transfers to real images before it's released. This role owns that measurement: the real test sets, the evaluation models, and the coverage, realism and privacy checks we report for each dataset.
What you'll work on
- Train evaluation models (detection, segmentation, classification or pose, depending on the domain) on our synthetic datasets and score them on held-out real images against a real-only baseline and synthetic-plus-real mixes, so each dataset comes with a measured result for how well it transfers.
- Build the real-image test sets those scores depend on, for each domain: sourcing, labeling, and splits that match the conditions a model will meet.
- Run ablations on the generation config to find which randomization ranges, materials, sensor settings and Cosmos-Transfer conditions move the real-world score and which only add frames.
- Define and maintain the dataset metrics: coverage of the parameter space and of rare cases, realism as the distance between synthetic and real feature distributions in a pretrained encoder such as DINOv2, drift between dataset versions, and privacy checks that no identifiable face, person or license plate from a real capture or scan ends up in a dataset.
- Audit exported labels against the scene. The hard cases are masks on transparent and reflective surfaces, instances that are almost fully occluded, slivers a few pixels wide, and boxes on objects cut by the frame edge. Check that our label conventions match the ones the real test sets use, such as amodal or visible boxes and what counts as one instance, since a mismatch looks like a transfer gap in the scores.
- Build the automatic label checks that run on every generated frame: small models and rules that flag empty, misaligned or implausible labels, written so they trace and export cleanly for inference inside the pipeline.
- Check Cosmos-Transfer output before it enters a dataset, including the cases where the transfer changes geometry that the labels still describe.
- Keep every result traceable to a dataset version, its scene parameters and seeds, and feed what you learn back into the generation configs yourself.
What you bring
- Detection, segmentation or classification models you have trained and debugged in PyTorch, starting from the failure cases when a result looks wrong.
- Careful evaluation: splits that don't leak, per-class and per-condition breakdowns, and confidence intervals that show how much a small real test set can tell you.
- Hands-on work with a gap between two image sources, such as synthetic and real, one camera and another, or day and night, and with ways to measure it.
- Fluency with labels and camera geometry: COCO and YOLO annotations, instance and semantic masks, depth maps, intrinsics and extrinsics.
- Python you would be fine handing to someone else, and experiments that rerun from their config a month later.
- Results written up so someone else can act on them: the plot, the config, what changed and what you would try next.
Good to have
- Synthetic data, domain randomization or sim-to-real work in robotics, driving or inspection.
- Generative-model metrics such as FID and KID, and a view on where they mislead.
- Pretrained backbones such as DINOv2, SAM or CLIP used as feature extractors or for label checks.
- Time spent with Cosmos or other video world models used for augmentation.
- Experience with medical imaging, automotive perception or damage assessment data.
- Dataset selection or active learning.
How we work
We're a small team based in Málaga, Spain, and remote-first. Communication is async, with deep focus time baked in, and there are no committees or layers of approval between your work and the product. We build on NVIDIA Omniverse, Replicator and Cosmos-Transfer, and we're members of the NVIDIA Inception Program.
How to apply
Write to us through the contact form with something you trained or evaluated and what the failure cases taught you. A repo, a notebook or a short write-up is enough. No cover letter.
or email hi@datadoo.ai