Audio AI Engineer
Job Description
Dev Technology is building a multilingual audio foundation to power a mobile translation experience that runs performantly on an iPhone. In this onsite role in Reston, VA, you will develop and improve a speech-to-text core, with ownership spanning multilingual ASR data, model adaptation for constrained devices, and reproducible evaluation of real-world quality.
What you’ll do
- Ingest, clean, segment, label, and version multilingual audio and transcript datasets, with specific attention to code-switching and borrowed-word behavior across the target language set.
- Fine-tune and compress large ASR models to meet iPhone-class memory, latency, and battery constraints while preserving transcription quality, using approaches such as LoRA/QLoRA, quantization, distillation, or other parameter-efficient and size-reduction techniques as needed.
- Create model packaging that supports selecting and downloading language-specific weights on demand based on operator context (for example, pulling only Chinese ASR weights for a Chinese-speaking interview).
- Build model behavior and evaluation around when to transcribe a borrowed English term as-is versus rendering it via transliteration or a native equivalent in the source language.
- Develop reproducible evaluation pipelines that measure word/character error rate, latency, and robustness to factors including accent, noise, speaking rate, and code-switching, and report results against defined success criteria per language and deployment target.
- Write clear model cards, dataset documentation, and evaluation write-ups so both technical and non-technical stakeholders can understand model purpose, comparisons, and known risks and limitations.
Requirements
- Bachelor’s degree in Computer Science, Data Science, Machine Learning, Computational Linguistics, or a closely related field.
- Strong data-engineering experience building production pipelines for large, messy, or unstructured audio and text datasets.
- Hands-on experience fine-tuning or adapting speech/audio models using parameter-efficient methods (LoRA, QLoRA, adapters) and/or compression techniques (quantization, distillation, pruning) for constrained hardware.
- Practical experience developing and evaluating ASR/speech-to-text models across multiple languages, including error analysis under real-world conditions such as accents, noise, and code-switching.
- Strong Python and SQL skills, with experience using tools such as PyTorch, Hugging Face Transformers/PEFT, torchaudio, and librosa.
- Experience deploying and monitoring production ML systems, including secure handling of sensitive audio, transcripts, and derived data in a regulated environment.
- Ability to explain model behavior, tradeoffs, and limitations to both technical and non-technical stakeholders.
Technologies
- Python, SQL, PyTorch
- Hugging Face Transformers, PEFT
- torchaudio, librosa
- LoRA, QLoRA, quantization, distillation, pruning
- iPhone, iOS, Swift, AVFoundation
Role scope
- This position owns the speech-to-text model: its data, training/adaptation, on-device size and latency, and transcription accuracy across languages.
- It does not own iOS application development, translation (source-language-to-target-language), or the Swift/AVFoundation integration layer; those are handled by a separate mobile engineering function, with close collaboration.
Salary and location
USD 80,000 - 160,000 per yearly
Reston, VA (onsite)
Benefits
- Generous and flexible time-off policy
- Flexible work schedules and telework options, including remote work availability for eligible projects
- Career development opportunities: mentorship program, Dev University technical and management training, hands-on learning through DevLab, tuition reimbursement, and paid training opportunities
- Industry-leading benefits including a choice of two health plans with dental and vision, flexible spending account, commuter benefits, life insurance, and more
- 401K matching with a 5% matching contribution
- Regular team and company social events including an annual party, happy hours, fitness challenges, and more
- Community engagement support, including employer match for donations and time off for volunteer efforts