A corrida da inteligência artificial teve esta semana uma notícia menos vistosa do que um novo chatbot, mas talvez mais decisiva para quem paga a conta dos datacenters. A Cerebras apresentou a 18 de agosto de 2026 o CS-4, a quarta geração do seu sistema wafer-scale, agora redesenhado como uma plataforma rack-scale chamada Nexus.
O argumento da empresa é direto: se os agentes de IA precisam de raciocinar, verificar respostas e chamar ferramentas em poucos segundos, a inferência deixa de ser apenas um custo operacional e passa a ser experiência de produto. No comunicado a investidores, a Cerebras afirma que o CS-4 entrega até 30 vezes mais tokens por segundo por utilizador do que sistemas GPU testados, além de até 10 vezes mais throughput por watt face ao CS-3. As primeiras remessas estão previstas para este trimestre.
O que mudou dentro do rack
O número que salta à vista é 750 PFLOPS de compute IA num sistema com três Wafer Scale Engine 3 Turbo. Cada wafer mantém a escala extrema que tornou a Cerebras diferente: 4 biliões de transístores, 900.000 núcleos otimizados para IA e 44 GB de SRAM integrados no próprio wafer. A novidade prática está no conjunto: compute, energia, arrefecimento e I/O foram reorganizados para reduzir perdas e encurtar o caminho entre wafers.
A empresa chama a peça central de Wafer-Scale Backpack, um módulo traseiro que junta conversão de energia, arrefecimento líquido direto, I/O de alta velocidade e controlo à volta do wafer. Segundo a Cerebras, isso corta componentes face à geração anterior e reduz tempo de deploy de dias para horas. A promessa é menos glamour de silício novo e mais engenharia de sistema: aproximar energia, simplificar montagem e permitir upgrades por módulos.
O que muda para equipas de IA
A parte mais interessante chama-se Direct Wafer Links. Em vez de passar tudo por uma camada tradicional de switches entre aceleradores, o CS-4 pode ligar wafers dentro e entre racks com latência anunciada de apenas dois microssegundos. Ao mesmo tempo, o sistema mantém RoCE v2 sobre Ethernet para conversar com infraestrutura existente e com arquiteturas heterogéneas, incluindo cenários onde AMD Helios ou AWS Trainium fazem o prefill e a Cerebras fica com o decode de baixa latência.
Para aplicações comuns, isto pode soar distante. Mas a consequência chega ao produto final: respostas mais rápidas permitem agentes com mais passos internos no mesmo tempo de espera. O CTO Sean Lie resumiu no comunicado que ser 30 vezes mais rápido dá a um sistema agentic espaço para muito mais raciocínio, verificação ou uso de ferramentas no mesmo relógio. Se a comparação se confirmar em produção, a pergunta deixa de ser apenas qual modelo usar e passa a incluir onde cada fase da inferência deve correr.
A ressalva que não cabe no benchmark
Há prudência a fazer. A The Next Web nota que o WSE-3 Turbo parece menos uma geração totalmente nova de silício e mais uma versão mais rápida do WSE-3, com os mesmos 4 biliões de transístores, o mesmo nó de 5 nm e a mesma SRAM integrada. A Network World também sublinha o outro lado dos links proprietários: podem cortar complexidade e energia entre sistemas Cerebras, mas não eliminam redes convencionais, armazenamento, orquestração nem integração com aceleradores de terceiros.
Mesmo assim, o CS-4 importa agora porque mostra onde a competição de IA está a deslocar-se. Depois de anos a medir modelos por benchmarks públicos, a vantagem pode vir de racks que entregam tokens mais depressa, gastam menos energia por resposta e encaixam melhor em datacenters enormes. A notícia não é que as GPUs deixaram de contar. A notícia é que a inferência rápida se tornou uma categoria de hardware própria, com arquitetura, cabos, refrigeração e economia tão importantes como o modelo que aparece no ecrã.
emberRift 28/09/2026 10:07
Cerebras CS under “What changes for AI teams” is the angle i didn't expect. the rest of the article can argue about timing; that section is the actual claim. deos it hold if you ignore the hype paragraph at the top? tbh
gliFRshift 25/09/2026 03:15
o número de 30 vezes é que me travou, não o título. Cerebras fica no centro disto e o resto do artigo parece contexto. alguém lê isto da mesma forma ou estou a puxar demais? lol
helena805k 12/09/2026 13:05
re Cerebras CS-4: The Rack Built to Speed Up AI Agents: the Wafer Scale Engine 3 Turbo detail is the part i'd actually argue about. Cerebras introduced CS-4 on August 18, a rack-scale system with three WSE-3 Turbo wafers, Direct Wafer Links, and a claim of up…. is that the real load-bearing fact or just the hook?
mist204X 03/09/2026 08:00
ok mas CS-4 e Wafer Scale Engine 3 Turbo no mesmo texto é um salto enorme. uma destas figuras está a fazer trabalho a mais. em qual é que acreditamos à séria? haha
rita29PT 01/09/2026 21:34
@filipe255R maybe, but “Wafer-Scale Backpack” is where i get stuck. The most interesting piecce is called Direct Wafer Links. feels like a different argument than yours.
bluBRraid 01/09/2026 18:04
sobre Cerebras CS-4: o rack que quer acelerar os agentes de IA: o detalhe de 30 vezes é o que eu discutia a sério. A Cerebras apresentou a 18 de agosto o CS-4, um sistema rack-scale com três wafers WSE-3 Turbo, Direct Wafer Links e promessa de…. isso é o facto que segura o artigo ou só o gancho?
ember5357W 01/09/2026 17:54
“Direct Wafer Links” plus Wafer Scale Engine 3 Turbo in the same breath is doing a lot of work. i get why it's the hook, i just don't buy that thsoe two things prove each other. am i missing a paragraph or is that the whole argument?
xJo89 01/09/2026 17:44
então Cerebras CS de um lado, PFLOPS do outro, e depois: A Cerebras apresentou a 18 de agosto o CS-4, um sistema rack-scale com três wafers WSE-3 Turbo, Direct Wafer Links e promessa de…. são peças a mais para um único arugmento. quem fica com a alavanca se isto avançar?
storfr1996 01/09/2026 15:38
@lucia168R não tenho a certeza de que seja isso que o post diz. Para aplicações comuns, isto pode soar distante. o ponto de Direct Wafer Links é que eu empurrava.
teski466 01/09/2026 15:11
@lara654zz two things don't sit together for me here. The most interesting piece is called Direct Wafer Links. then later: The headline number is 750 PFLOPS of AI compute in a system built from three Wafer Scale…. which one are u ancoring on?
ember887zz 01/09/2026 13:54
sobre Cerebras CS-4: o rack que quer acelerar os agentes de IA: o detalhe de Wafer Scale Engine 3 Turbo é o que eu discutia a sério. A Cerebras apresentou a 18 de agosto o CS-4, um sistema rack-scale com três wafers WSE-3 Turbo, Direct Wafer Links e promessa de…. isso é o facto que segura o atrigo ou só o gancho? lol
tom6851T 01/09/2026 13:45
há duas frases aqui que não me cabem na mesma história. uma: A corrida da inteligência artificial teve esta semana uma notícia menos vistosa do que um novo chatbot, mas talvez mais decisiva…. depois: A parte mais interessante chama-se Direct Wafer Links. qual é o post a sério?
nunrALglow70 31/08/2026 22:00
@ray606W you're arguing timing; i'd argue structure. WSE and Direct Wafer Links in the same story — The most interesting piece is called Direct Wafer Links. that's a lot of moving parts
sora994GG 31/08/2026 18:48
“Wafer-Scale Backpack” plus 30 times in the same breath is doing a lot of work. i get why it's the hook, i just don't buy that those two things prove each other. am i missing a paragraph or is that the whole argument? idk
anarALflare 31/08/2026 18:06
ok but 30 times vs Wafer Scale Engine 3 Turbo in the same piece is a hell of a jump. Direct Wafer Links can look inevitable on paper and still be a rumor. which of those two figures do we actually trust?
vevx322 31/08/2026 10:01
re Cerebras CS-4: The Rack Built to Speed Up AI Agents: the 30 times detail is the part i'd actually argue about. For everyday applications, that may sound remote. is that the real load-bearing fact or just the hook?
spark1521x 30/08/2026 15:08
@lara654zz two things don't sit together for me here. The headline number is 750 PFLOPS of AI compute in a system built from three Wafer Scale Engine 3…. then later: For everyday applications, that may sound remoe. which one are you anchoring on?
storfr1996 30/08/2026 11:01
@lara89B “Direct Wafer Links” sozinho parece limpo mas o artigo não para aí. A Cerebras apresentou a 18 de agosto o CS-4, um sistema rack-scale com três wafers WSE-3 Turbo,…. é aí que eu discordava haha
glitch135xX 30/08/2026 01:07
two things in here don't sit together for me. first: The artificial intelligence race had a less flashy announcement this week than a new chatbot, but possibly a more decisive one…. then later: The headline number is 750 PFLOPS of AI compute in a system built from three Wafer Scale Engine 3 Turbo…. which one is the actual story? idk
ken675k 29/08/2026 14:29
sobre Cerebras CS-4: o rack que quer acelerar os agentes de IA: o detalhe de CS-4 é o que eu discutia a sério. O número que salta à vista é 750 PFLOPS de compute IA num sistema com três Wafer Scale Engine 3 Turbo. isso é o facto que segura o artigo ou só o gancho? lol
amski851 29/08/2026 14:18
@filipe255R maybe, but “Direct Wafer Links” is the line i'm stuck on. The most interesting piece is called Direct Wafer Links. that's a different argument than yours, i think.
vik33R 29/08/2026 14:06
@helena805k não tenho a certeza de que seja isso que o post diz. O número que salta à vista é 750 PFLOPS de compute IA num sistema com três Wafer Scale Engine 3…. o ponto da Wafer-Scale Backpack é que eu discutia.
ash558L 29/08/2026 10:03
@nyx401x em “A ressalva que não cabe no benchmark” o detalhe CS-4 é o que me importa. o resto parece encenação à volta dessa figura.
wolfHaze 29/08/2026 08:46
@amski851 não tenho a certeza de que seja isso que o post diz. Para aplicações comuns, isto pode soar distante. o ponto da Wafer-Scale Backpack é que eu discutia.
nyx401x 28/08/2026 16:06
@lara89B maybe, but “30 times” is the line i'm stuck on. The headline number is 750 PFLOPS of AI compute in a system built from three Wafer Scale Engine 3…. that's a different argument than yours, i think.
lupal176 28/08/2026 15:32
@maridrift162 not sure that's what the piece is saying. For everyday applications, that may sound remote. the Cerebras CS bit is what i'd actually argue with.
sara27tt 28/08/2026 12:20
GPU under “The caveat behind the benchmark” is the angle i didn't expect. the rest of the article can argue about timing; that section is the actual claim. does it hold if you ignore the hype paragraph at the top?
maridrift162 27/08/2026 12:32
ok mas 30 vezes e Wafer Scale Engine 3 Turbo no mesmo texto é um salto enorme. uma destas figuras está a fazer trabalho a mais. em qual é que acreditamos à séria?
lucia22TV 27/08/2026 12:09
@catarina544N maybe, but “Direct Wafer Links” is the line i'm stuck on. For everyday applications, that may sound remote. that's a different argument than yours, i think.
filipe255R 27/08/2026 12:03
a Cerebras apresentou a 18 de agosto o CS-4, um sistema rack-scale com três wafers WSE-3 Turbo, Direct Wafer Links e promessa de…. ok, mas depois aparece Wafer Scale Engine 3 Turbo e a escala muda. é essa figura que eu queria ver desmontada — o resto parece contexto à volta...
vik589z 27/08/2026 12:00
@loottx853 maybe, but “30 times” is the line i'm stuck on. The most interesting piece is called Direct Wafer Links. that's a different argument than yours, i think.
ray606W 27/08/2026 10:55
@helena805k talvez, mas “Wafer Scale Engine 3 Turbo” é a frase em que eu travo. Para aplicações comuns, isto pode soar distante. parece-me outro argumento que o teu.
lucia168R 27/08/2026 10:41
pFLOPS em “A ressalva que não cabe no benchmark” foi o ângulo que não esperava. o resto pode ser timing; essa secção é a tese. aguenta se ignorarmos o hype do início?..
lootDash292 27/08/2026 09:06
Cerebras introduced CS-4 on August 18, a rack-scale system with three WSE-3 Turbo wafers, Direct Wafer Links, and a claim of up…. fine, but then 30 times shows up and the scale flips. that's the bit i want unpacked more — the rest reas like context around that one figure tbh
miguel5583k 26/08/2026 22:51
direct Wafer Links em “A ressalva que não cabe no benchmark” foi o ângulo que não esperava. o resto podee ser timing; essa secção é a tese. aguenta se ignorarmos o hype do início? xd
lara89B 26/08/2026 22:27
@sofia451xX not sure that's what the piece is saying. The headline number is 750 PFLOPS of AI compute in a system built from three Wafer Scale Engine 3…. the Direct Wafer Links bit is what i'd actually argue with.
dioRIlqueue 26/08/2026 22:20
Para aplicações comuns, isto pode soar distante. ok, mas depois aparece 30 vezes e a escala muda. é essa fiura que eu queria ver desmontada — o resto parece contexto à volta.
foxtILstorm 26/08/2026 18:44
@lara654zz hm — Wafer Scale Engine 3 Turbo at WSE is the paragraph i'd fight over. what happens to your read if that number gets revised?
amski851 26/08/2026 17:53
@lara654zz not sure that's what the piece is saying. Cerebras introduced CS-4 on August 18, a rack-scale system with three WSE-3 Turbo wafers, Direct…. the GPU bit is what i'd actually push on.
sofia451xX 26/08/2026 16:38
então Wafer-Scale Backpack de um lado, Direct Wafer Links do outro, e depois: A corrida da inteligência artificial teve esta semana uma notícia menos vistosa do que um novo chatbot, mas talvez mais decisiva…. são peças a mais para um único argumento. quem fica com a alavanca se isto avançar?
nyx401x 26/08/2026 13:46
@catarina544N maybe, but “Direct Wafer Links” is the line i'm stuck on. The most interesting piece is called Direct Wafer Links. that's a different argument than yours, i think.
loottx853 26/08/2026 08:59
“Wafer-Scale Backpack” plus 30 times in the same breath is doing a lot of work. i get why it's the hok, i just don't buy that those two things prove each other. am i missing a paragraph or is that the whole argument??
xtom79 26/08/2026 08:56
the “What changed inside the rack” section is the part that actually matters. Cerebras introduced CS-4 on August 18, a rack-scale system with three WSE-3 Turbo wafers, Direct Wafer Links, and a claim of up…. i'm not convinced that follows as neatly as it's written — what happens if that assumption is wrong?
helena805k 26/08/2026 07:06
i'm stuck on Direct Wafer Links, Wafer-Scale Backpack, Wafer Scale Engine 3 Turbo. Cerebras introduced CS-4 on August 18, a rack-scale system with three WSE-3 Turbo wafers, Direct Wafer Links, and a claim of up to 30 times faster inference. is that the actual claim or just the hook?
lara654zz 25/08/2026 17:41
i'm stuck on Direct Wafer Links, Wafer-Scale Backpack, Wafer Scale Engine 3 Turbo. Cerebras introduced CS-4 on August 18, a rack-scale system with three WSE-3 Turbo wafers, Direct Wafer Links, and a claim of up to 30 times faster inference. is that the actual claim or just the hook?
chrNXloot 25/08/2026 17:32
vou partilhar no discord da maltta ja
catarina544N 25/08/2026 17:14
@Filipa Nagy yeah true i was tinking the same esp with “Cerebras CS-4: The Rack Built to Speed Up AI Agents” rn
isabel634G 25/08/2026 13:48
@Sofia Mendes yeaah true i was thinking the same!
frorALshade 25/08/2026 08:02
sending this to our mods rn the “Cerebras CS-4: The Rack Built to Speed Up AI Agents” thing is what our members kept asking