Desert Ant Labs logo

ML Researcher

Desert Ant Labs

Full-time
5-10 years experience
Remote OK

Key Skills

Machine Learning
Speech And Audio Modeling
Computer Vision
Video Modeling
Multilingual Text Modeling
Model Compression
PyTorch
Python
Model Evaluation
Dataset Development
Synthetic Data Generation
Quantization
Pruning
Knowledge Distillation
On-Device Deployment
Benchmarking

Job Description

Train our speech, vision, video, and text models, make them smaller and faster on real devices, and ship them in Detail and Subwave, our own apps. About Desert Ant Labs Desert Ant Labs is an on-device AI lab in Europe. We're building efficient frontier intelligence, bottom up. We make the fastest model for each task, built to run on the billions of phones, laptops, and browsers people already own. Every model is lightning fast and private, and has no per-call cost. Little brains in every product. With our native SDKs for Swift, Kotlin, and JavaScript, you add a model to an app in a few lines of code. Data stays with your customer, and the feature works without a cloud provider. Every role here pushes the limits of what a phone or a laptop can do. About the role You work on our speech, vision, video, and text models with the research team. We hire researchers at every level, senior included. Each researcher goes deep in one area: speech and audio, video and vision, multilingual text, or efficiency. Most work spans areas: labeling who speaks uses audio and video, picking clips starts from a transcript, and every model has to fit on a phone. Senior researchers take a model from the product question to the release in an app, then the next one. The goal is models on a phone that people assume need a server. Our models today Speech and audio: Voz transcribes speech in 25 languages, Clear cleans up a phone recording to studio sound, Uhm finds and removes filler words, Ear identifies which of 99 languages is spoken, and Align gives each word a timestamp. Voz is NVIDIA's Parakeet, which we convert and compress. Video and vision: Clips picks the best short clips from a longer recording, working from its transcript, in 100 languages. Eye, Face, and Who, which labels who is speaking from the audio and the video, are in beta. Text: Redact filters PII in 27 languages, Gist tags topics in 101 languages, Tongue detects 84 languages from three words, and Title suggests a title and description for any text. Models for structured extraction and hate speech detection are in beta. Efficiency: every model is quantized, distilled, or pruned to fit, converted for each platform, and benchmarked on real devices. Shrinking a model from 8MB to 2MB can take weeks, and the work is wasted when the 8MB model already fits the product, so we work on the limits that matter first. Next: getting the beta models out of beta, and new models in every area. How we do research Our researchers train models, and they also profile, convert, and benchmark those models on phones and laptops. A model's speed and memory use on a phone decide whether a product can use the model. We hire people with experience in both research and engineering: researchers who care about implementation and speed, and engineers who moved into training. We pick each model's default settings for the products that use the model, such as which kinds of personal data Redact removes when the developer changes nothing, or whether Uhm leaves a filler word in or risks cutting a real word. You make those decisions with the team, and we expect you to say so when you disagree. We want a published benchmark for every language a model supports, and the same quality in each. We aren't there yet for every model, and closing those gaps is part of the work. What you will do Train and improve models, from the dataset to the release in an app. Build evaluations, and publish the benchmarks with every release. Generate synthetic and augmented data, and test on real recordings, videos, and text from the devices and rooms people use. Quantize, prune, and distill models to fit the size, latency, and memory limits of phones, laptops, and the browser, and keep as much of each model as possible on the Neural Engine, NPU, or GPU, with the runtime engineers. Measure latency, memory, heat, and battery use on real hardware before you report a number. Build models that combine audio, video, and text, such as labeling who speaks when. Decide with the Detail and Subwave teams what a model should do, then measure the model in the app. Write the model card and the release post for each model. You might be a fit if you Have trained and shipped a speech, vision, video, or text model under a size or latency budget, on a phone, a single-board computer, or in a browser. We value a shipped product more than a paper. Go deep in one of speech and audio, video and vision, multilingual text, or model compression, and want to work across the others. Have worked on every step from data to a model running on a phone or in a browser. Can design an evaluation and defend its results in public, including in languages other than English. Write solid Python and know PyTorch well. For a senior role: have shipped models for at least 5 years, and can walk through one, from what it was for to how you measured it and what you'd change. Write clean, tested code, and spot a weak change in review, whether a person or an agent wrote it. Plan and run your work through coding agents such as Claude Code, with a low tolerance for slop. Bonus points Worked on speech enhancement, diarization, forced alignment, or streaming speech models. Built efficient temporal models for video, or face or speaker models under privacy constraints. Worked on low-resource languages, or fine-tuned a small language model for one task. Know an on-device runtime in depth: Core ML, LiteRT, ONNX Runtime, ExecuTorch, or MLX. Wrote Metal, CUDA, or NPU kernels, or contributed to ExecuTorch, MLX, llama.cpp, or coremltools. How we work We're a small, flat team, and we build the models, the SDKs, and the apps that use them. Detail and Subwave, our own apps, run our models in production. Start what needs starting without waiting to be asked, and finish what you start. Take on work outside your role when a project needs you. We ship quickly, so we cut scope until only the part users notice is left. Anyone can comment on your work or redo your draft, and we say early when work isn't ready. We read that feedback as help. We judge the work by what shipped and what changed because of it. Nobody counts hours. What we offer Salary and equity, based on level. In our Amsterdam office, or remote anywhere between Eastern Time in North America (UTC-5) and Central European Time (UTC+1), so everyone shares part of the working day. Recruiters and agencies We hire directly. We don't reply to emails from recruiters or agencies, and we don't want their outreach about this role or any other. Location: Amsterdam or Remote (UTC-5 to UTC+1) Apply on our site: https://desertant.com/jobs/ml-researcher/

Core Responsibilities

Train and improve speech, vision, video, and text models from dataset development through release in the company’s apps, including building evaluations and publishing benchmarks. Optimize models for size, latency, memory, heat, and battery use on real devices, and collaborate with product and runtime teams to measure and ship them.

Requirements

Candidates should have trained and shipped a speech, vision, video, or text model under size or latency constraints, with experience spanning data preparation through on-device or browser deployment. Strong Python and PyTorch skills, the ability to design and defend evaluations, and clean, tested code are expected; senior candidates should have at least five years of model-shipping experience.

Benefits

  • Salary
  • Equity

About Desert Ant Labs

Industry: Technology, Information and Internet

Company size: 2-10 employees

Desert Ant Labs is a European frontier AI lab building opinionated on-device audio and visual models. We put little brains in every product to unlock new experiences, with no inference cost and no data leaving your device. We're building intelligence for developers, on the 6 billion devices people already own.

Added Today