Persona AI Inc is seeking a Staff-level Robotics ML/Data Engineer to architect and scale multimodal robotics data pipelines. The role focuses on transforming raw, in-the-wild egocentric video and dense sensor streams into high-fidelity training assets that directly support foundation model development.
Responsibilities
- Design cross-modal validation systems that verify agreement between video, proprioception, force/haptic signals, and language annotations. Examples include reprojecting robot state into the image plane to confirm video-state consistency, and using VLM-assisted checks to assess whether instructions align with observed behavior.
- Orchestrate hand-tracking, segmentation, depth estimation, 3D reconstruction, and pose-tracking components. Retarget human demonstrations into robot trajectories, and run simulation-in-the-loop validation using kinematic feasibility, physics replay, and motion-consistency filtering to ensure synthesized data is physically grounded rather than only visually plausible.
- Implement robust data augmentation methods to expand expert trajectories and improve learning robustness, including spatial transformations, temporal scaling, synthetic viewpoints, and sensor noise injection.
- Create unified state-action representations across different embodiments, coordinate frames, rotation conventions, gripper and hand parameterizations, and sampling rates. Apply per-dimension validity masking and per-source normalization so that onboarding a new robot or sensor becomes a configuration update instead of a code rewrite.
- Build dataset tooling that enables researchers to query, visualize, and audit data, including clip browsers, trajectory viewers, and annotation review user interfaces. Translate model-failure analysis into new curation rules and targeted re-collection requests.
- Architect end-to-end ingestion pipelines that convert raw, unstructured sources into indexed, queryable, training-ready datasets. Include temporal segmentation of long recordings into action clips, metadata and scene-graph extraction, embedding-based retrieval, and language annotation workflows.
Requirements
- M.S. or Ph.D. in Computer Science, Data Engineering, Machine Learning, Robotics, or a related field.
- Deep expertise in Python with extensive experience using PyTorch, especially for custom dataloaders for multimodal datasets.
- Experience analyzing and processing complex time-series data from force-torque sensors, load cells, or tactile arrays, with careful alignment to visual frames.
- Strong knowledge of video processing pipelines and libraries including OpenCV, FFmpeg, and Decord, including managing I/O bottlenecks for terabyte-scale video datasets.
- Solid robotics and 3D geometry knowledge covering coordinate frames and transforms, rotation representations, camera intrinsics and extrinsics, forward/inverse kinematics, and URDF.
- Proven ability to implement programmatic and generative data augmentation for computer vision and time-series data.
Technologies
Python, PyTorch, OpenCV, FFmpeg, Decord, URDF, Ray, Apache Spark, Open X-Embodiment, DROID, AgiBot World, EgoDex, SAM-family, MANO, SMPL, Omniverse, MuJoCo, NVIDIA robotic software stack, NVIDIA's robotic software stack, VLM
Benefits
- Competitive compensation
- Performance-based bonus
- 99% employer covered medical benefits
- Early-stage equity
- Competitive PTO
- Company-wide paid winter break between December 24th and January 2nd
- Full access to advanced tools
Bonus Skills
- Experience with NVIDIA’s robotic software stack (Open X-Embodiment, DROID, AgiBot World, EgoDex, or similar).
- Comfort using modern perception tools as a user, including segmentation (SAM-family), monocular depth, hand/body pose estimation (MANO/SMPL), and 6-DoF object pose tracking and point tracking. Experience integrating and evaluating these components in a pipeline is expected.
- Familiarity with distributed data processing systems such as Ray and Apache Spark.
- Background in generating or utilizing synthetic robotic data through simulation using Omniverse and MuJoCo.
- Experience integrating spatial awareness or tactile data representations (for example, Fourier encoding) into visual pipelines.
Job Details
- Department: Software
- Reports To: Teleoperations Lead
- Employment Type: Full-Time
- Location: Houston, TX (onsite); Houston, TX or Pensacola, FL