Software Engineer at Hirenza
Washington, District of Columbia, 20022, USA
Listed on 2026-10-01
-
Software Development
AI Engineer (Applied/Software), Machine Learning/ ML Engineer
Leo Technologies is a pioneering company dedicated to developing innovative solutions that leverage advanced audio and speech processing technologies. Committed to transforming how organizations extract meaningful insights from audio data, Leo Technologies specializes in building scalable, high-precision transcription and speech recognition systems. Their focus on cutting‑edge machine learning models, cloud infrastructure, and real‑time processing enables them to serve diverse industries including security, media, healthcare, and enterprise analytics.
With a culture rooted in innovation, collaboration, and continuous improvement, Leo Technologies strives to push the boundaries of what is possible in speech and audio intelligence, ensuring their clients stay ahead in a rapidly evolving digital landscape.
The Role
We are seeking a highly skilled Transcription Engineer to join our Platform Team in a remote, work-from-home capacity. This role is central to our mission of extracting actionable intelligence from complex audio environments. As a Transcription Engineer, you will be responsible for enhancing transcription quality across various workflows by experimenting with Automatic Speech Recognition (ASR) models, integrating third‑party services, and developing tooling that guarantees accuracy, reliability, and scalability.
The ideal candidate possesses a robust software engineering background, with expertise in Python, audio processing, and applied machine learning techniques tailored for speech recognition. This position combines hands‑on engineering, data science, and Dev Ops skills, involving experimentation with ASR models and deploying services into production offers a challenging yet rewarding opportunity for professionals passionate about audio, language, and building high‑quality systems that power real‑world intelligence use cases.
The ideal candidate should have a strong foundation in software engineering, with a minimum of five years of professional experience in speech processing, NLP, or transcription systems. Proficiency in Python is essential, along with familiarity with system‑level programming when necessary. Experience with ASR frameworks such as Whisper, Kaldi, Vosk, or NVIDIA NeMo is highly desirable. Candidates should have a solid understanding of audio engineering tools like ffmpeg and Sox, and techniques for denoising and voice enhancement.
Knowledge of speaker diarization, speaker recognition, and multi‑language ASR challenges is important. Experience with data analysis tools such as Pandas, Num Py, and Jupyter for evaluating model performance is required. A good understanding of cloud deployment and Dev Ops practices, including Docker, Kubernetes, and serverless architectures, is also necessary. The ability to work independently in a fast‑paced environment, make tradeoffs, and deliver results with minimal supervision is crucial.
Bonus points are awarded for experience in fine‑tuning ASR models on domain‑specific datasets, real‑time streaming pipelines, search and retrieval systems like Elasticsearch, and prior work in audio forensics or noisy‑channel speech analysis.
The core responsibilities of this role include leading efforts to improve transcription quality by evaluating, testing, and fine‑tuning various ASR models, whether commercial APIs or open‑source solutions. You will build and optimize pipelines capable of speaker identification, diarization, multi‑language support, and noise‑robust transcription in challenging audio conditions. Developing and maintaining resilient services that integrate multiple ASR providers to ensure flexible and reliable workflows is essential.
Collaborating with platform engineers to…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).