• 80/20 AI
  • Posts
  • Introducing a New Standard for AI Hardware

Introducing a New Standard for AI Hardware

Audit Your AI Stack Before the September 1 Deadline

Advertise here | 6-min Read

Chatbots and voice agents are so 2024. Now you can build video agents that look just like us that can work for you.

Meet PAL Maker by Tavus, the no-code way to spin up a video agent that sees you, hears you, and talks back face-to-face in real time.

Send your PALs to a Google Meet as a notetaker, talk to one like your teammate, build one as an onboarding guide for your customers

Upload your own face or pick out of 100 stock options. PALs respond in over 45+ languages, and have memory that carries across conversations.

If you like poking at the edge of what AI can do, consider this your next rabbit hole. Tavus already powers 200k+ builders and the world's most forward-thinking companies

Introducing a New Standard for AI Hardware

WHAT’S HAPPENING AI TODAY

1. COLM 2026: LLMs cost up to 1,431x more than embeddings at equal quality — the benchmark that reframes your AI cost model: (cite index="16-1">COLM 2026: LLMs cost up to 1,431x more than embeddings at equal quality. The finding is from the Conference on Language Modeling 2026 and covers specific task categories where embedding models — smaller, faster, cheaper — match or exceed frontier LLM quality. The 1,431x cost difference is not an average; it is the top end of the range on tasks where embeddings are well-suited.

2. SemiAnalysis published its Jalapeño deep dive — taped out in 16 months, 13.4 PFLOPs, 700W vs NVIDIA Rubin's 900–1,150W: (cite index="15-1">SemiAnalysis published a deep dive on OpenAI's first custom inference chip Jalapeño, taped out with Broadcom in just 16 months on TSMC N3P. The B0 stepping hits 13.4 PFLOPs of MXFP4 at 700W versus Rubin's 900–1,150W, pairs HBM4 at 15.4TB/s, and posts 700+ tokens per second per user on DeepSeek R1 and around 1,400 tokens per second on shorter contexts.

3. NVIDIA server prices rising 15%+ from 2027 — Vera Rubin and Grace Blackwell configurations affected: (cite index="15-1">Nvidia's contract server builders have told Microsoft, Google, and Oracle that prices on AI server systems will rise more than 15% starting on shipments in early 2027, hitting flagship Vera Rubin and Grace Blackwell configurations. The price increase compounds on top of Samsung's 15% foundry price increase and SK Hynix's HBM capacity constraints. For any team modelling AI infrastructure costs beyond 2026.

AI NEWS HIGHLIGHT

COLM 2026: LLMs cost 1,431x more than embeddings at equal quality on specific tasks — the most important cost benchmark of the year. September 1 makes deploying frontier LLMs for embedding-suited tasks materially more expensive to get wrong.
Jalapeño deep dive: 13.4 PFLOPs at 700W, 1,400 tok/s on short context, 16-month tape-out — 700W vs NVIDIA Rubin's 900–1,150W. Power efficiency compounds at data centre scale. OpenAI used AI to design the chip.
NVIDIA server prices +15% from 2027 — Vera Rubin and Grace Blackwell both affected — hardware floor rising on three fronts simultaneously: NVIDIA servers, Samsung foundry, HBM capacity. Per-token prices falling; hardware costs climbing.
AM Intelligence ordered 9,000 Vera Rubin systems for $8B — largest single AI hardware order of August — Greenko Group's Hyderabad-based data centre. India's AI infrastructure build is now at $8B single-order scale.
South Korea's Wrtn Technologies raised $72M Series C at $722M+ valuation — AI content platform, 1 trillion won valuation. K-AI is its own market. Wrtn is the largest consumer AI company in Korea by valuation.
Intel Xeon 7 Diamond Rapids slipped to 2027 — 256 cores, drops hyperthreading entirely for HPC — AMX with FP8, AVX 10.2, 128 PCIe 6.0 lanes. Originally 2026. Every major chip program this year has slipped at least once.
Claude Sonnet 5: 4 days left at $2/$10 — September 1 is $3/$15. The COLM 2026 1,431x finding is the most important context for making that routing decision correctly.

Someone You Know Is Looking for Their Next Role.

Recommend trusted nursing job opportunities through Jobstream™ and help another nurse take the next step in their career. Earn when they apply through your referral.

Audit Your AI Stack Before the September 1 Deadline

Prompt: COLM 2026 found LLMs cost up to 1,431x more than embedding models at equal quality on specific tasks. Sonnet 5 moves from $2/$10 to $3/$15 in 4 days.

Our AI stack: [list every model, what task it handles, and monthly spend].

1. Classify each workload — retrieval, classification, similarity, or generation. The first three are where embeddings match frontier quality at a fraction of the cost. Which of our LLM workloads are actually embedding tasks in disguise?

2. Calculate the saving — for each embedding candidate: cost at $3/$15 vs text-embedding-3-large at $0.13/M tokens. Show the monthly delta.

3. Give me the decision — for every workload: embed, stay on Sonnet 5, or move to a cheaper frontier model. No open items.

End with the one migration that delivers the highest saving for the least engineering effort. That is the one to do this weekend.

TOP TRENDING AI TOOLS

That’s a Wrap

SPONSOR US

Get your business in front of over 90k+ AI professionals

8020AI is the world’s #1 AI Newsletter, Read by 90k+ professionals from leading companies such as Google, OpenAI, Meta, and Microsoft.

We've assisted in promoting Over 500 AI-Related ProductsWill yours be the next?

What We Can Offer:

  • Launch an Advertising Campaign

  • Introduce New Product or Features

  • Other Business Cooperation

Or Email our founder Alamin at [email protected]

FEEDBACK

How was your experience with 8020AI today?

How was 8020AI today?

Login or Subscribe to participate in polls.

Login or Subscribe to participate in polls.

If you have specific feedback or anything interesting you’d like to share, please let us know by replying to this email.