A dataset run assembles USD scenes, renders them, writes images and labels to disk, and runs models over the output: Cosmos-Transfer passes, automatic label checks, and the evaluation models that score the result. A slow step costs GPU hours whether it runs on the GPU or leaves one idle, and GPU hours set how many variants we can try before a dataset is good. This role finds where the time goes and takes it out, rewriting hot paths in C++ and CUDA, and makes the model steps fast and reproducible by tracing and exporting them for inference runtimes.
What you'll work on
- Profile full dataset runs with Nsight Systems, NVTX ranges, Tracy, perf and py-spy, and find out whether a run is bound by scene assembly, rendering, annotator readback, encoding and disk I/O, or Python on Kit's main thread.
- Keep a benchmark of fixed scenes and seeds that runs on every pipeline change. It flags drops in throughput and rises in host or GPU memory, and checks that a regenerated run matches the original frame by frame, with a record of what can't be bit-exact (path-tracer sampling, denoisers, multi-GPU rendering).
- Move hot scene-assembly and randomization code from Python to C++: Sdf-level authoring inside an
SdfChangeBlockwhen the Usd API is too slow, Fabric and USDRT for per-frame changes so a randomization step doesn't recompose the stage, and instancing, payloads, load rules and population masks that cut stage load time. - Write Omniverse Kit extensions and OmniGraph nodes in C++ for the parts of Replicator on the hot path, keep annotator output on the GPU, and overlap readback and file writing with rendering.
- Cut startup and first-frame time on render machines (shader and MDL compilation caches, texture loading), and keep VRAM inside budget on large scenes.
- Write CUDA or Warp kernels for the per-frame post-processing that shows up in the profile: run-length encoding of masks, boxes from instance IDs, depth conversion and resizing, with JPEG encoding on the GPU through nvJPEG or nvImageCodec.
- Trace and export the models that run inside the pipeline, starting with the label checks that run on every frame and the evaluation models:
torch.exportto ONNX or Torch-TensorRT, and TorchScript tracing where an older model still needs it. Check each exported model against eager PyTorch, comparing tensors within a tolerance and the label decisions themselves at FP16 or INT8. Keep the built engines and timing caches, and pin driver, CUDA, TensorRT and weight versions per GPU type, so the same input gives the same labels next month. - Speed up Cosmos-Transfer runs: context-parallel inference across GPUs, BF16 or FP8 and fewer denoising steps where the output still passes our checks, running only the control branches a dataset uses, and scheduling so no GPU sits idle between clips.
- Make the data loaders for evaluation-model training and pipeline inference keep the GPUs busy: GPU image decoding, sharded datasets, pinned memory and fewer host-device copies.
What you bring
- Modern C++ (C++17 or later) shipped in a codebase where speed mattered, written with allocation, data layout, cache behavior and threading in mind from the start.
- Something you made measurably faster with a profiler, and the ability to show where the time went before and after.
- CUDA kernels you have written and tuned with Nsight Compute, and the judgment to see when the kernel is fine and the copy or the sync is the problem.
- PyTorch models you have exported with
torch.export, TorchScript tracing or ONNX, including rewriting the ops and data-dependent control flow that wouldn't export and setting dynamic batch and resolution, and then run under TensorRT or ONNX Runtime with outputs checked against the original. - Concurrency in practice: thread pools and task systems such as TBB, lock contention, async I/O and CUDA streams.
- C++ bound into Python with pybind11 or nanobind, since Python stays the glue around the pipeline.
- Comfort in a large C++ codebase you didn't write, on Linux: CMake, sanitizers, gdb, and reading the library source when the docs run out.
Good to have
- The OpenUSD C++ API, Hydra, or Omniverse Kit and Carbonite extension development.
- NVIDIA Warp, or Fabric and USDRT inside Kit.
- Renderer or game engine internals: GPU frame timing, Vulkan, OptiX or other ray tracing work.
- TensorRT plugins or custom CUDA ops for PyTorch.
- Inference work on diffusion or video models: attention kernels, quantization, FP8 and mixed precision.
- GPU image pipelines with nvJPEG, NVENC or DALI.
How we work
We're a small team based in Málaga, Spain, and remote-first. Communication is async, with deep focus time baked in, and there are no committees or layers of approval between your work and the product. We build on NVIDIA Omniverse, Replicator and Cosmos-Transfer, and we're members of the NVIDIA Inception Program.
How to apply
Write to us through the contact form with something you profiled and made faster: what the profile showed, what you changed and what it did. A repo, a PR or a short write-up is enough. No cover letter.
or email hi@datadoo.ai