Senior ML Engineer (Token Factory)

Company hidden · Remote · Czechia, Europe, Germany, Israel, Netherlands, UK ·
Fully remoteFull-timeSenior5+ years
COMPENSATION
Salary available to subscribers
0 applicants so far

Listed on Jobicy

  • Senior ML Engineer
  • Fully remote (EU/Global)
  • Python + PyTorch
  • LLM Inference Optimization
  • Large-scale GPU Cloud

Join the Token Factory team at the company, where you will build high-performance inference and fine-tuning platforms for foundation models. Your work will involve pushing hardware limits to optimize cost-per-token and latency across a massive GPU cloud infrastructure.

You will tackle complex challenges in inference optimization, engine support, and low-precision training. The team values deep technical expertise in transformer architectures and GPU performance tuning to drive production speedups for various LLM architectures.

What you'll do

  • Identify LLM inference bottlenecks to drive production speedups at scale
  • Implement novel speculative decoding architectures for diverse LLM designs
  • Optimize components of dense and MoE models for autoregressive or parallel execution

Also in this posting

  • What you'll need
  • Nice to have
  • Stack
  • How you'll work
  • What you get
Application link is subscriber-only

Unlock the direct hiring contact and apply link for this role, plus every other job in the database.