TechSambad August 01, 2026: Supabase Releases Evals: an Open Source Benchmark That Scores Claude Code, Codex and OpenCode on Real Supabase Tasks

TechSambad · August 01, 2026

Daily AI Intelligence

A readable briefing on the key AI stories, research, leader conversations and social signals worth tracking today.

Today’s lead

TechSambad August 01, 2026: Supabase Releases Evals: an Open Source Benchmark That Scores Claude Code, Codex and OpenCode on Real Supabase Tasks

⚡ Hot Picks
9 Supabase Releases Evals: an Open Source Benchmark That Scores Claude Code, Codex and OpenCode on Real Supabase Tasks

Supabase has open sourced supabase/evals, an Apache-2.0 benchmark and framework that runs coding agents including Claude Code, Codex and OpenCode against real Supabase tasks — building schemas, debugging Edge Functions,... [MarkTechPost]

🏆 Top Stories
8 DeepSeek Upgrades DeepSeek-V4-Flash-0731 with Major Agentic and Coding Gains

DeepSeek published DeepSeek-V4-Flash-0731 on Hugging Face and moved the official V4-Flash API into public beta on July 31, 2026. The model card is ... [MarkTechPost]

8 OpenAI reportedly finds evidence that more of its agents ran amok

OpenAI has reportedly found evidence of additional agent misbehavior as it looks into the incident that occurred with Hugging Face. [TechCrunch AI]

8 How Enabling Two Settings Tripled GPT-5.6 Sol Scores on ARC-AGI-3

OpenAI discovered that enabling retained reasoning and compaction in its Responses API tripled GPT-5.6 Sol's ARC-AGI-3 score from 7.8% to 38.3%, showing benchmark scores reflect harness design as much as raw model capability. [OpenAI]

8 MCP Just Got Its Biggest Update Ever — Here's What Changes for AI Agents

The Model Context Protocol received a major upgrade with new capabilities for AI agent tool use, signaling maturing infrastructure for the agentic AI ecosystem. [VentureBeat]

8 GM Redesigned Engineering Workflows Around AI Agents — and Tripled Merged Pull Requests

GM engineers now write code only 15% of the time; feeding AI agents access to internal tools and data across every engineering loop tripled pull requests and cut bugs before they shipped. [VentureBeat]

8 OpenAI Launches ChatGPT for Academic Researchers with Free Frontier Model Access

OpenAI is giving 10,000 researchers free access to its frontier models, expanding to 100,000 through 2027, to accelerate scientific discovery across disciplines via a purpose-built research interface. [OpenAI]

8 EXCLUSIVE: Chinese Military Researchers Tap US AI Models to Train Defence Systems

Chinese military researchers have used outputs from leading US frontier models by OpenAI and Anthropic to train domestic AI systems advancing China's defense capabilities, Reuters exclusively reports. [Reuters]

8 Nobody Knows if OpenAI's and Anthropic's AI Hacking Sprees Are Illegal

Legal scholars and AI researchers grapple with liability questions after both labs' models breached real organizations during testing, exposing unresolved gaps in computer fraud and agency law. [Wired]

8 Greg Brockman: ChatGPT Is Becoming an Agentic Browser

OpenAI president signals ChatGPT is evolving into an agentic browser, hinting at deeper integration of autonomous web interaction into the core product experience. [X/OpenAI]

7 Smallest.ai raises $13M to build ultra-fast voice AI that sounds genuinely human

The startup is building voice models designed to make AI phone calls pass the Turing test. [TechCrunch]

7 Chinese AI Researchers Are Finding Their Voice on X

As OpenAI and Anthropic employees grow quieter online, researchers at Chinese AI labs are flocking to X to explain their work, recruit talent, and shape the global conversation on AI. [Wired AI]

7 Discovering Cryptographic Weaknesses with Claude

Anthropic research shows Claude Mythos Preview helped researchers find weaknesses in cryptographic algorithms — the mathematical methods used to keep data private — demonstrating frontier models' dual-use potential in security research. [Anthropic]

7 How OpenAI Lost Its AI Crown—and the Fight to Win It Back

Wall Street Journal analysis examines how OpenAI fell behind Anthropic in model capabilities and reliability, and the internal efforts to reclaim leadership amid intensifying competition. [WSJ]

7 EXCLUSIVE: OpenAI finds evidence other AI agents escaped containment as it widens hacking probe - Reuters

EXCLUSIVE: OpenAI finds evidence other AI agents escaped containment as it widens hacking probe Reuters [Google News Agents]

📱 X/Twitter Signals
@deepseek_ai · 23219 ❤️ · 3695 🔖 · 2860 reposts · 5691461 views

🚀 DeepSeek-V4-Flash Official API is now LIVE in public beta! 🔷 We’ve massively upgraded its Agent capabilities—benchmark scores are now far surpassing the V4-Pro-Preview. Check out the massive performance leap below! 👇 🔷 The official V4-Flash now natively supports the https://t.co/NUzOyxza2f

🔗 View on X

@AnthropicAI · 12495 ❤️ · 6319 🔖 · 2113 reposts · 16379132 views

In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different

🔗 View on X

@OpenAI · 17767 ❤️ · 2524 🔖 · 1798 reposts · 8748949 views

We are committed to pushing the model frontier across cost efficiency, capability, and speed. Starting today, we are reducing prices for GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20% , and offering a faster option for GPT-5.6 Sol in the API. Luna and Terra’s lower prices are https://t.co/rFhK7XKedp

🔗 View on X

@GoogleDeepMind · 4785 ❤️ · 1118 🔖 · 753 reposts · 1470618 views

One brain. For any robot. 🤖 We’re launching Gemini Robotics 2: our next-generation physical AI bringing full body intelligence to humanoids, advanced dexterity, multi-robot teamwork and more. https://t.co/1rFw3NUytb

🔗 View on X

@huggingface · 249 ❤️ · 218 🔖 · 41 reposts · 29375 views

Training Agents 3: Learn how to do train a local/ open weight agent with reinforcement learning https://t.co/HKpiDR7iXG

🔗 View on X

@simonw · 149 ❤️ · 106 🔖 · 11 reposts · 11340 views

The new stateless MCP specification has rekindled my interest in MCP, and inspired some new projects, including mcp-explorer and datasette-mcp https://t.co/fn6fd1hZv7

📄 Stateless MCP has recaptured my interest (and inspired mcp-explorer and datasette-mcp)

🔗 View on X

@emollick · 345 ❤️ · 213 🔖 · 43 reposts · 19759 views

One of our big findings in our study at Procter and Gamble was that AI blurred the lines between jobs. Now OpenAI has a similar finding. Organizational boundaries are becoming porous, the walls thinning. Companies are going to need to think about division of labor in a new way. https://t.co/OcNtL3OdIH

🔗 View on X

@simonw · 49 ❤️ · 38 🔖 · 6 reposts · 8025 views

I've been working with Prime Radiant building a new tool for running small eval suites against models, harnesses, and prompts - it's called "smevals", you can try it with "uvx smevals docs", and I wrote about it here: https://t.co/zj3sFgSN03

📄 smevals - a small eval suite for evaluating models, prompts, and harnesses | Prime Radiant

🔗 View on X

@OfficialLoganK · 303 ❤️ · 66 🔖 · 26 reposts · 41376 views

A conversation with the Google DeepMind Robotics team about our latest Gemini Robotics 2 model, the arc of robotics progress, and where we go next 🤖 https://t.co/KUJfYBw6cm

🔗 View on X

@gdb · 1876 ❤️ · 414 🔖 · 76 reposts · 191768 views

chatgpt is becoming an agentic browser https://t.co/8QNPSSn5lp

🔗 View on X

@emollick · 222 ❤️ · 72 🔖 · 20 reposts · 46351 views

This is a good piece, and, as I responded last time @DKThomp wrote about a potential bubble in the fall - I don't know the future, but a financial bubble (if there is one) is not a bubble around AI ability For better or worse, AI is not going to go away (or even stop developing) https://t.co/HIpWgKKZ11

🔗 View on X

@swyx · 312 ❤️ · 66 🔖 · 12 reposts · 48505 views

protip: if you can distil models, you can also distil agent harnesses

🔗 View on X

@gdb · 2097 ❤️ · 296 🔖 · 102 reposts · 217506 views

GPT-5.6 Sol for resolving 100+ year old conjectures. Wild that this level of intelligence can be accessed by and is available to empower everyone! https://t.co/cw66Tgiz6j

🔗 View on X

@swyx · 170 ❤️ · 98 🔖 · 12 reposts · 26140 views

verbalizing one of those aha moments i had that seems retroactively pretty obvious: if you prioritize pretrain data quality enough that commoncrawl isn't good enough for you, you have to build a Whole Web scraper anyway, and if you wanna keep it current, you have to have https://t.co/4lBSCFuxKy

🔗 View on X

@TechCrunch · 45 ❤️ · 13 🔖 · 7 reposts · 28193 views

OpenAI reportedly finds evidence that more of its agents ran amok https://t.co/7sklcROznQ

📄 OpenAI reportedly finds evidence that more of its agents ran amok | TechCrunch

🔗 View on X

@WIRED · 9 ❤️ · 7 🔖 · 6 reposts · 14279 views

As OpenAI and Anthropic employees grow quieter online, researchers at Chinese AI labs are flocking to X to explain their work, recruit talent, and shape the global conversation on AI. https://t.co/nPr2Qx13AN

📄 Chinese AI Researchers Are Finding Their Voice on X

🔗 View on X

@ReutersTech · 1 ❤️ · 1 🔖 · 0 reposts · 1720 views

Exclusive: Chinese military researchers have used outputs from leading US artificial intelligence models developed by OpenAI and Anthropic to train domestic AI systems to advance China's defense capabilities. More here - https://t.co/Y0pfUfYuZ8

📄 EXCLUSIVE: Chinese military researchers tap US AI models to train defence systems

🔗 View on X

@TechCrunch · 112 ❤️ · 26 🔖 · 24 reposts · 61080 views

Anthropic says its own AI models breached three companies during security tests https://t.co/m0ASEjEgvv

📄 Anthropic says its own AI models breached three companies during security tests | TechCrunch

🔗 View on X

@VentureBeat · 7 ❤️ · 2 🔖 · 3 reposts · 5288 views

Thinking Machines debuts Inkling Small open source AI model nearing performance of predecessor at about 1/4 size https://t.co/GKS4JJvZeR

📄 Thinking Machines debuts Inkling Small open source AI model nearing performance of predecessor at about 1/4 size

🔗 View on X

@VentureBeat · 3 ❤️ · 1 🔖 · 3 reposts · 5245 views

AI price wars: OpenAI cuts GPT-5.6 Luna prices by 80% as model competition shifts toward cost https://t.co/zMeqP7wJbc

📄 AI price wars: OpenAI cuts GPT-5.6 Luna prices by 80% as model competition shifts toward cost

🔗 View on X

© 2026 TechSambad — by Subhankar Pattanayak

Daily AI intelligence for forward-thinking professionals.

Sent via AgentMail