Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

KV caching status?

What's the point of 1000tok/s if you have to do prefill on every agentic turn which at 100k depth would make it 1.5 min latency every turn?



Information about RAM type/size and connection topology of the RAM to be used for context cache seems to be conspicuously absent from the slick looking marketing materials.


There's a few more details at the bottom of this page: https://www.cerebras.ai/blog/introducing-cerebras-cs-4

44GB on-chip-sram * 3 chips. Per chip: 43.2 PB/s memory access + 53.5 PB/s on-chip fabric bandwidth + 2.4 Tbits/s "IO" bandwidth (I think that means their RoCE v2 RDMA over Ethernet interface).

I suspect there might be a certain amount of customization for how much RAM they attach when you order it.


They have managed to make the external link 300GBps/2us. Cs3 was 150/5.

This is 1/3rd blackwells nvlink c2c bandwidth already. Not too bad. We can make KV cache offload work with that I suppose.

If magically KV cache was not an issue, pipeline parallelism on cerebras can be quite pleasant. As for the KV cache offload, I have hopes their CPO solution they're trying with that canadian company ends up bearing fruit.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: