ProductHardwareLibraryCareersDocs
Book a call

The research lab focused on inference.

TrainingAutomatic fine-tuning from your traffic.InferenceOpen-model inference, pay based on your latency requirement.
NotesFrontier notes on inference, systems and hardware.MethodsHow we operate, and why.

Latest writing

  • Announcing Carat: An inference engine designed for Gemma 4Sep 2, 2026
  • ThesisJul 13, 2026
  • Affordable and open intelligenceSep 9, 2026
See all writing
A Gradiated inference package, shown from above.

Runs models.Not bills.

The chip design is based on your requirements. Our first chip is based on efficiency, not speed. It's designed for specific open source models, to suit ultra intense, long running cloud workloads at the lowest cost imaginable.

We are hiring our founding silicon and hardware engineers. We'd love to talk to you if this is something you could build.

Build with us
Announcing Carat: An inference engine designed for Gemma 4Designing Carat, an inference engine built specifically around the model architecture of Gemma 4.Sep 2, 2026
ThesisMost of the car is missing.Jul 13, 2026
Affordable and open intelligenceMore intelligence for less energy, open to everyone.Sep 9, 2026

Product

  • Training
  • Inference
  • Service tiers
  • Models and pricing

Developers

  • Quickstart
  • API reference
  • Errors
  • Status

Company

  • Thesis
  • Careers
  • Library

Legal

  • Privacy
  • Terms
  • Acceptable use
  • DPA
  • Security

© 2026 Gradiated Ltd

Cambridge, --:--:-- GMT