Senior ML Engineer (Token Factory)
Company hidden · Remote · Czechia, Europe, Germany, Israel, Netherlands, UK ·
COMPENSATION
Salary available to subscribersListed on Jobicy
- Senior ML Engineer
- Fully remote (EU/Global)
- Python + PyTorch
- LLM Inference Optimization
- Large-scale GPU Cloud
Join the Token Factory team at the company, where you will build high-performance inference and fine-tuning platforms for foundation models. Your work will involve pushing hardware limits to optimize cost-per-token and latency across a massive GPU cloud infrastructure.
You will tackle complex challenges in inference optimization, engine support, and low-precision training. The team values deep technical expertise in transformer architectures and GPU performance tuning to drive production speedups for various LLM architectures.
What you'll do
- Identify LLM inference bottlenecks to drive production speedups at scale
- Implement novel speculative decoding architectures for diverse LLM designs
- Optimize components of dense and MoE models for autoregressive or parallel execution
Also in this posting
- What you'll need
- Nice to have
- Stack
- How you'll work
- What you get
Application link is subscriber-only
Unlock the direct hiring contact and apply link for this role, plus every other job in the database.