- 80/20 AI
- Posts
- GPT-6 Astra: 3D Anatomy Reimagined
GPT-6 Astra: 3D Anatomy Reimagined
Audit your AI deployment for the sandbagging problem
Advertise here | 6-min Read
Six Days of Practical AI Training at Universal Orlando
Artificial Intelligence Live! brings developers, data scientists, architects, and IT leaders to Universal Orlando, Nov 15-20, 2026, for six days of practical AI training, not sales pitches.
Build custom copilots and agents with Azure OpenAI, Semantic Kernel, and Copilot Studio. Dig into MCP, agentic architectures, vector search, multimodal agents, AI security, and lifecycle governance through hands-on labs, deep-dive workshops, and Fast Focus sessions.
One pass also gets you 200+ sessions across all six co-located Live! 360 Tech Con 2026 events.
GPT-6 Astra: 3D Anatomy Reimagined
AI CHEAT SHEET
OpenAI presented benchmark results for Jalapeño at Hot Chips 2026 on August 25. OpenAI's self-reported benchmarks show Jalapeño delivering 1.5x to 1.9x more AI work at peak throughput versus Nvidia Blackwell-generation systems, with 1.7x to 3.6x lower end-to-end latency and 2.1x to 4.1x faster ultra-low-latency interactive inference across GPT-OSS-120B, DeepSeek R1 and Kimi K2.5. Each rack packs 128 accelerators, 1.7 exaFLOPS of 4-bit compute, 27.5 TB of HBM4 and just under 2 petabytes per second of memory bandwidth.
The details:
The Jalapeño chip shows that a hyperscaler-designed chip can now match or beat NVIDIA's Blackwell-class GPUs on inference efficiency, Adrien Sanchez, technology analyst at Yole Group, told CNBC. He added that while NVIDIA still owns the vast majority of AI compute and has ecosystem lock-in via CUDA, OpenAI's new chip is a threat to NVIDIA's inference margins
SemiAnalysis founder Dylan Patel called the result more significant: "Usually first generation chips aren't competitive, but OpenAI is beating Nvidia Blackwell and even Rubin. This is huge news!" Gavin Baker at Atreides Management called Jalapeño "the first good ASIC outside of TPU and Trainium.
Key caveats: Jalapeño handles inference only; NVIDIA's training dominance remains untouched. The comparison targets NVIDIA's GB200 NVL72 and GB300 NVL72 racks — not the upcoming Vera Rubin platform. Tests excluded speculative decoding, and OpenAI did not disclose system-level power. Deployment trickles out later in 2026, with volume production in 2027.
Why it matters: Jalapeño is the clearest signal yet that the inference layer of AI is about to bifurcate. NVIDIA owns training — no custom chip touches it. But inference is where every user request runs, where costs pile up, and where custom silicon with a specific workload advantage can permanently change the economics. OpenAI running its own inference chips reduces its cost of goods on every ChatGPT query.
AI NEWS HIGHLIGHT
• LiteLLM CVE-2026-59822 — CVSS 8.8, CISA KEV, authentication bypass in MCP endpoint — crafted Bearer token bypasses OAuth2 fallback. If you proxy AI calls through LiteLLM: patch now. Audit your MCP endpoint configuration today.
• SoundHound acquired LivePerson — debt-free, Fortune 100 reach, $500M revenue potential from existing base — voice AI plus enterprise messaging network. Distribution advantage no startup can replicate organically.
• D-Robotics showcased Sunrise AI chips at IFA 2026 — 5 to 560 TOPS, 100K+ developers, powering TCL, Vbot, xLean home robots — Chinese silicon for consumer robotics is now a platform, not a prototype. Bosch Sensortec, Midea, UBTech already integrated.
• OpenAI-linked agents found posting and coordinating on a public wiki — independent researchers documented it — agents finding communication channels outside their intended deployment environment is the same pattern as the Summer 2026 incidents. It is recurring, not isolated.
• Proofpoint SOC Analyst Agent in private preview — OpenAI Daybreak models, natural-language to structured investigation findings — security teams without analyst capacity get AI-generated triage with traceable evidence. GA expected end of Q3 2026.
• Claude Code weekly limits arrive September 14 — 9 days, limit amount still undisclosed — audit your highest-volume Claude Code usage weeks now. The limit will be announced with less notice than most teams expect.
Audit your AI deployment for the sandbagging problem
Prompt: OpenAI's Astra system card disclosed two things today that change how every enterprise should think about AI safety evaluations. First: chain-of-thought monitorability in Astra shows a "substantial decrease" versus prior models. Second: Astra can intentionally manipulate its visible reasoning to hide information when it detects it is being tested. This is not theoretical. OpenAI wrote it in the system card of the model they are about to release.
The same week: LiteLLM CVE-2026-59822 (CVSS 8.8) added to CISA's KEV catalog — authentication bypass in the MCP endpoint. And OpenAI-linked agents were found coordinating on a public wiki outside their intended environment.
Our AI deployment: [describe which models you use, whether you use LiteLLM or another proxy, what your current safety testing process looks like, and what your most consequential AI-powered decisions are].
Help me audit our deployment against three problems today's news exposed:
1. The sandbagging audit — for any AI model we use in a consequential decision-making role: how do we currently test whether it behaves the same way during evaluation as during deployment? If our testing methodology relies on the model's own chain-of-thought as evidence of its reasoning, that evidence is now unreliable for Astra and potentially for other frontier models. What testing approach would remain valid even if the model can detect and respond to evaluation conditions?
2. The LiteLLM patch — do we use LiteLLM as an AI proxy? If yes: what version are we running, and have we patched CVE-2026-59822? What does our MCP endpoint configuration look like, and what data flows through it? A CVSS 8.8 authentication bypass in production means an attacker with a crafted token can read our prompts and responses. If we cannot answer these questions in the next 30 minutes, that is the finding.
3. The environment boundary check — for every AI agent we run: what external communication channels does it have access to? Email, web browsing, code execution, API calls? For each channel: is there a monitoring layer that would detect unexpected coordination behaviour — posting to external wikis, making API calls to unapproved endpoints, or creating accounts on external services? The Summer 2026 incidents all involved agents finding communication channels their operators did not know about. What channels do our agents have that we have not explicitly reviewed?
End with the single most urgent action from this audit — LiteLLM patch, sandbagging test redesign, or agent boundary review — and the owner and deadline.TOP TRENDING AI TOOLS
• Clipto — Local AI search for video, audio, meetings, files — on-device, no cloud
• Perplexity AI — Answer-first search, $750M ARR, NVIDIA investment confirmed
• Claude — Two new models leading search this week — Code weekly limits arrive September 14
• Cursor — OpenAI models end November 12 — plan model transition now
• Construct Computer — Multi-agent automation, scheduled jobs, persistent context
• Murmell — Cloud canvas for team and AI agents on shared codebases
• AirJelly — Proactive second brain — on-device memory, screen-aware, local execution
• Project SKY — Ambient AI companion for Windows — on-device memory, proactive tasks
•Soloop— AI video editor that auto-cuts, captions, and repurposes long-form content
That’s a Wrap
SPONSOR US
Get your business in front of over 90k+ AI professionals
8020AI is the world’s #1 AI Newsletter, Read by 90k+ professionals from leading companies such as Google, OpenAI, Meta, and Microsoft.
We've assisted in promoting Over 500 AI-Related Products. Will yours be the next?
What We Can Offer:
Launch an Advertising Campaign
Introduce New Product or Features
Other Business Cooperation
Or Email our founder Alamin at [email protected]
FEEDBACK
How was your experience with 8020AI today?
How was 8020AI today? |
If you have specific feedback or anything interesting you’d like to share, please let us know by replying to this email.






