×
Register Here to Apply for Jobs or Post Jobs. X

Applied Scientist, Agent Evaluation & Adaptive Model Routing

Job in Austin, Travis County, Texas, 78701, USA
Listing for: Bitdeer Technologies Group
Full Time position
Listed on 2026-09-03
Job specializations:
  • IT/Tech
    Machine Learning/ ML Engineer, AI Engineer (Applied/Software), AI Business & Operations
Job Description & How to Apply Below

Job Title

About Bitdeer:

Bitdeer is a world-leading technology company for Bitcoin mining and AI cloud. Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers. Apart from designing industry-leading ASIC chips and manufacturing mining rigs, the Group handles complex processes involved in computing across the value chain. This includes equipment procurement, transport logistics, datacenter design and construction, equipment management, and network and facility operations.

Bitdeer also offers advanced cloud capabilities to customers with a high demand for artificial intelligence. Headquartered in Singapore, Bitdeer operates globally with a diversified 3 GW energy portfolio, and deploys Bitcoin mining and HPC datacenters in the United States, Bhutan, Norway, Canada, Malaysia, and Ethiopia.

About Bitdeer AI Lab:

Bitdeer AI Lab is a frontier AI lab under Bitdeer, a global-leading computing power solutions provider. Guided by long-termism, we are committed to exploring the frontiers of artificial intelligence with the ambition, courage, and determination to build technologies that can truly change the world.

Our mission is to turn energy into intelligence that people can actually afford to use. Inference is where that happens: every product built on a model is bounded by what it costs to run, so the economics of serving decide what gets built  work on this from the ground up, from the power and datacenters we own to the software that turns them into tokens — and we continue to invest in and expand the infrastructure behind it

What You Will Be Responsible For
  • This role builds the evaluation and decision systems that make agentic inference measurable, reliable, and adaptive. You will own and extend our LLM and agent evaluation pipeline, develop representative task suites, and measure behavior at the trajectory level across task success, tool use, quality, cost, latency, and token consumption. You will research and prototype adaptive model-routing strategies inside a shared agent harness, including rule-based baselines, cascades, stage-aware routing, uncertainty-aware selection, and escalation and recovery policies.

    You will take promising methods from offline evaluation through traffic replay, shadow testing, and internal pilots, while working closely with the MaaS and platform engineering teams on production integration. The role owns evaluation methodology, routing policy, and research prototypes; production gateway reliability, billing, access control, and SLA engineering are collaborative responsibilities with the MaaS and platform teams
How You Will Stand Out
  • Bachelor's, Master's, or PhD in Computer Science, Machine Learning, Statistics, Electrical Engineering, or a related field, with substantial hands-on experience in LLM evaluation, agentic systems, applied machine learning, or adaptive inference
  • Strong programming ability in Python and practical experience with PyTorch and modern data and evaluation tooling; able to independently build reliable experimental pipelines and internal research prototypes
  • Demonstrated experience designing or operating evaluation pipelines for LLMs or agents, including task-level and trajectory-level metrics, dataset construction, automated scoring, regression testing, and failure analysis
  • In addition to hands-on LLM or agent evaluation experience, candidates should have implementation-level depth in at least one core area: model selection and routing, uncertainty estimation and calibration, cascading and escalation, or stage-aware agent inference. Experience with preference modeling, contextual bandits, online learning, or broader adaptive inference methods is a plus.
  • Hands-on experience evaluating multi-turn or tool-using agents, including task completion, tool-call correctness, planning failures, recovery behavior, and cost and latency trade-offs
  • Rigorous experimental practice — controlled comparisons, honest baselines, statistical analysis, and the ability to explain clearly what an evaluation result does and does not prove
  • Ability to analyze per-request and per-trajectory quality, cost, latency, and token usage, and construct…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary