EngineerJobs.io
← Back to all jobs

Job Description

Persona AI Inc is seeking a Staff-level Robotics ML/Data Engineer to architect and scale multimodal robotics data pipelines. The role focuses on transforming raw, in-the-wild egocentric video and dense sensor streams into high-fidelity training assets that directly support foundation model development.

Responsibilities

  • Design cross-modal validation systems that verify agreement between video, proprioception, force/haptic signals, and language annotations. Examples include reprojecting robot state into the image plane to confirm video-state consistency, and using VLM-assisted checks to assess whether instructions align with observed behavior.
  • Orchestrate hand-tracking, segmentation, depth estimation, 3D reconstruction, and pose-tracking components. Retarget human demonstrations into robot trajectories, and run simulation-in-the-loop validation using kinematic feasibility, physics replay, and motion-consistency filtering to ensure synthesized data is physically grounded rather than only visually plausible.
  • Implement robust data augmentation methods to expand expert trajectories and improve learning robustness, including spatial transformations, temporal scaling, synthetic viewpoints, and sensor noise injection.
  • Create unified state-action representations across different embodiments, coordinate frames, rotation conventions, gripper and hand parameterizations, and sampling rates. Apply per-dimension validity masking and per-source normalization so that onboarding a new robot or sensor becomes a configuration update instead of a code rewrite.
  • Build dataset tooling that enables researchers to query, visualize, and audit data, including clip browsers, trajectory viewers, and annotation review user interfaces. Translate model-failure analysis into new curation rules and targeted re-collection requests.
  • Architect end-to-end ingestion pipelines that convert raw, unstructured sources into indexed, queryable, training-ready datasets. Include temporal segmentation of long recordings into action clips, metadata and scene-graph extraction, embedding-based retrieval, and language annotation workflows.

Requirements

  • M.S. or Ph.D. in Computer Science, Data Engineering, Machine Learning, Robotics, or a related field.
  • Deep expertise in Python with extensive experience using PyTorch, especially for custom dataloaders for multimodal datasets.
  • Experience analyzing and processing complex time-series data from force-torque sensors, load cells, or tactile arrays, with careful alignment to visual frames.
  • Strong knowledge of video processing pipelines and libraries including OpenCV, FFmpeg, and Decord, including managing I/O bottlenecks for terabyte-scale video datasets.
  • Solid robotics and 3D geometry knowledge covering coordinate frames and transforms, rotation representations, camera intrinsics and extrinsics, forward/inverse kinematics, and URDF.
  • Proven ability to implement programmatic and generative data augmentation for computer vision and time-series data.

Technologies

Python, PyTorch, OpenCV, FFmpeg, Decord, URDF, Ray, Apache Spark, Open X-Embodiment, DROID, AgiBot World, EgoDex, SAM-family, MANO, SMPL, Omniverse, MuJoCo, NVIDIA robotic software stack, NVIDIA's robotic software stack, VLM

Benefits

  • Competitive compensation
  • Performance-based bonus
  • 99% employer covered medical benefits
  • Early-stage equity
  • Competitive PTO
  • Company-wide paid winter break between December 24th and January 2nd
  • Full access to advanced tools

Bonus Skills

  • Experience with NVIDIA’s robotic software stack (Open X-Embodiment, DROID, AgiBot World, EgoDex, or similar).
  • Comfort using modern perception tools as a user, including segmentation (SAM-family), monocular depth, hand/body pose estimation (MANO/SMPL), and 6-DoF object pose tracking and point tracking. Experience integrating and evaluating these components in a pipeline is expected.
  • Familiarity with distributed data processing systems such as Ray and Apache Spark.
  • Background in generating or utilizing synthetic robotic data through simulation using Omniverse and MuJoCo.
  • Experience integrating spatial awareness or tactile data representations (for example, Fourier encoding) into visual pipelines.

Job Details

  • Department: Software
  • Reports To: Teleoperations Lead
  • Employment Type: Full-Time
  • Location: Houston, TX (onsite); Houston, TX or Pensacola, FL

Similar Jobs