Sequencing and depth
Start in software. Go down the stack.
Start with the useful product and go down the stack: harness, inference software, infrastructure, then silicon. Each layer earns the right to build the next and makes the whole system better.
There is an overwhelmingly large range of things we could work on simultaneously. So much so, that it could cause us to fail. We think a key to winning is keeping things boringly-simple.
We are starting in software whilst renting GPUs. We will directly sell our API to developers and will offload excess capacity to inference routers. This achieves two things. First, it means we are building extraordinarily clear requirements for our hardware. Second, we can get to significant scale, so the economics of a chip are worth it.
We will then ship a chip, as we start designing our own data centers. We are likely to ship multiple chips over time, focusing on efficiency over latency. This enables us to get to more scale, enabling us to justify data centers, and enabling us to purchase chips from those more focused on ultra low latency use cases (which we consider to have more technical risk).


