Inference Research
Founding Inference Research Engineer
Build and optimize the complete inference stack for open-weight models on existing hardware.
- Location
- Cambridge, London
- Office
- Monday, Wednesday, Friday
- Compensation
- £240,000–£500,000 base + significant equity
About Gradiated
Gradiated is a research lab focused on inference. We build software and hardware for open-weight models.
Our goal is to maximize tokens per joule. We work across model behavior, software, systems, and silicon to find better ways to run models.
We treat each model as a mathematical function. We map that function onto hardware and build a custom software stack for it.
We are early. You will help set the technical direction and the way the company works.
Why this role exists
The best inference system does not come from one isolated optimization. Model structure, numerical formats, kernels, memory, scheduling, and serving affect each other.
You will work across these layers. You will turn research ideas into measured systems that run real models.
What you will do
- Study open-weight models from their tensors and numerical behavior through to production execution.
- Build important parts of the inference stack from first principles.
- Write and optimize CUDA or HIP kernels for exact model shapes.
- Optimize and measure the stack on current NVIDIA and AMD hardware.
- Design model representations, quantization methods, memory layouts, and cache policies.
- Build execution, scheduling, distributed inference, and serving paths.
- Measure speed, energy use, model quality, correctness, and reliability on real workloads.
- Contribute useful work to open-source projects when that is the best path.
- Share clear evidence. Stop ideas that do not survive measurement.
Problems you may work on
- Model-specific kernels and fused execution paths.
- Low-bit formats that preserve model quality.
- KV-cache layout, movement, compression, and reuse.
- Sparse and mixture-of-experts execution.
- Speculative and parallel decoding.
- Prefill and decode scheduling across several GPUs.
- Performance models that explain the gap between hardware limits and measured results.
- New runtime designs for long-running agent workloads.
What we are looking for
- You have built a difficult ML or systems project that you can explain in detail.
- You can work in Python and a systems language such as C++ or Rust.
- You can write CUDA or HIP kernels, or you can learn a new GPU programming model quickly.
- You understand modern model execution, including attention, memory movement, and numerical precision.
- You use profiling and experiments to find the real limit.
- You can work without a complete specification.
- You write clearly and expose your reasoning to review.
This role is not for you if
- You want to own one narrow layer and hand work to another team.
- You only want to connect existing inference frameworks.
- You only want to run experiments in notebooks.
- You want research papers to have priority over working systems.
- You optimize benchmark numbers without checking quality and correctness.
- You need a fixed roadmap before you can start useful work.
How we work
We prefer evidence to assertion. State the question, measure the result, and record what you learned. A failed idea is useful when it gives us a clear decision.
We work in the office on Monday, Wednesday, Friday. We value working in person to solve hard problems, but we are flexible. Some of the team work abroad for short periods.
Pay and equity
The base salary is £240,000–£500,000. Every offer also includes significant equity.
We pay at the top of the market because talent density is what will allow us to build the most efficient inference. We want a small team of exceptional people who can solve hard problems together.
Interview process
- A 60-minute call about your goals and the role.
- We get together in person and talk through the problems we will work on.
- A paid in-person working session, so we can get to know you and you can meet the team.
- References and an offer decision.
Apply
Let's talk.
If this work interests you, we would like to hear from you. You do not need to prepare anything. We will start with a conversation.
Start a conversation →