Sr. Research Engineer
Listed on 2026-08-05
-
IT/Tech
Data Annotation/ AI Labeling, AI Evaluation
The Opportunity
Adobe's Sound Design AI group (SODA) is looking for a driven Data/ML engineer to push the boundaries of audio GenAI. Join the team behind Firefly Generate Sound Effects and multiple AI models that have shipped in Adobe products. We're a small, collaborative and efficient research team looking for highly motivated candidates of all levels with a passion for audio and video data.
If the prospect of productizing cutting-edge research into impactful tools for Adobe's creative users sounds exciting, this is it.
We're hiring a data lead to own the training data behind our generative audio models end to end. Our models learn from large-scale audio and video corpora, and the quality, balance, and integrity of that data is one of the biggest levers on model performance. This person is the single owner of what's in the data, why, and how we know it's good.
Responsibilities- Build, deploy, and monitor large-scale data pipelines for ingesting, filtering and preprocessing audio and video at scale, with reproducibility and versioning.
- Run in-house models for inference over millions of audio/video files, and build data exploration tools to support this.
- Curate and select training data using both quantitative metrics and qualitative judgment. Maintain an ongoing understanding of the corpus: distribution, balance, diversity, coverage, and redundancy.
- Train or finetune models on pipeline outputs, evaluate their behavior, and use those findings to drive new data experiments and ablations. Scale training and optimize throughput.
- Drive data licensing and acquisition: identify gaps and opportunities, define briefs, and work with vendors and producers to license and curate new datasets.
- Own how training labels and annotations are produced, validated, and improved, including model-assisted labeling at scale.
- Pipeline-building experience at scale for audio and/or video data with reproducibility and versioning
- Deep audio domain knowledge: common datasets, quality metrics and their failure modes, codecs, normalization, and filtering. Strong "ears" and critical listening ability. Video knowledge a plus.
- Generative-ML research experience to make data decisions, design and run training experiments independently
- Strong command of evaluation in audio and video modelling.
- Data acquisition and licensing experience, including writing briefs and communicating across stakeholders.
- Excellent communication skills
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).