
This Week in All Things AI covers key developments in models, agents, tools, infrastructure, and policy curated from discussions in the All Things AI Telegram group.
If you follow AI for work, research, investing, or just to understand where the technology is heading, this weekly brief is a concise way to scan the most important launches, risks, and resources in a few focused minutes.
The week of 26th July to 1st August 2026 was defined by a decisive shift from raw model capability toward the economics, reliability, and operational design of agentic AI. Microsoft introduced its MAI-Cyber-1-Flash cybersecurity model and Project Perception, while benchmark work comparing Kimi K3 across Claude Code, Hermes, and Kimi Code showed that the harness can drive order-of-magnitude differences in token consumption even when task success rates remain similar. OpenAI also cut GPT-5.6 pricing, and discussion around cheaper Chinese models, region-specific coding plans, and hybrid model stacks reinforced the view that “good enough” intelligence is becoming a major competitive advantage.
Agent infrastructure, privacy, and open deployment were equally prominent. MCP’s move to a stateless core makes remote servers substantially easier to place behind load balancers and on serverless or edge infrastructure, while new search, crawling, and browser APIs from providers including Firecrawl, Octen, Context.dev, and TinyFish highlighted the growing importance of high-quality web access for agents. At the model layer, Fish Audio, Superwhisper and Cohere pushed more capable local speech tooling; MiniMax released its open-weight H3 multimodal video model; DeepSeek substantially upgraded V4-Flash’s agentic coding performance; and Poolside launched a cross-platform workspace for supervising parallel coding agents. Meanwhile, the indexing of supposedly private AI conversations and a public call by frontier-lab employees to deliberately pace AI development kept security and governance firmly in view.
The sections that follow walk through these items day by day, with short context and links so you can dive deeper into the pieces most relevant to your work or interests.
via the Relentless podcast [quite underrated at only 25K subscribers] TI Morse with Sam Altman
Sam drops some mind-blowing insights on how AI is completely rewiring the rules for startups, arguing that we are literally living in the "singularity" right now. He dives deep into why we need to trust in exponential AI progress, how to thrive as a leader in absolute chaos, and why superintelligence won't actually bring us the 4-hour workweek.
There's also some incredible behind-the-scenes intel on OpenAI's ruthless prioritization—like deciding to pause Sora to focus on coding agents—and why massive investments in chips, energy, and robotics are the real bottlenecks to our future.
Kevin Kelly highlights how AI tokenomics are shifting user behavior toward lower-cost, "good enough" models as expenses become a key factor, potentially making leading-edge AI a commercial liability.
The linked WSJ article reports US companies adopting cheaper Chinese open-weight models for routine tasks, using hybrids where premium models plan and affordable ones execute, achieving major savings like reducing daily AI agent costs from high figures to $100 per agent.
This reflects a broader "thrift-maxxing" trend in AI, prioritizing cost efficiency and model flexibility over maximum power, similar to how mass-market computing products historically outsold specialized high-end hardware.
via Yat Siu
Animoca Research launched "The Agentic Age," its first monthly brief on the agentic economy, emphasizing that abundant AI coverage requires sharper judgment for capital allocators, operators, and builders
Microsoft introduces MAI-Cyber-1-Flash, an AI model trained for cybersecurity, and launches Perception, an agentic security system to patch vulnerabilities. It says MAI-Cyber-1-Flash and MDASH, its vulnerability identification harness, deliver “world-class performance at 50% of the cost of leading models”
via Alex
Anthropic forgot to add noindex and basic privacy blocks to shared conversations, making hundreds of thousands of privately shared conversations (including on enterprise plans) publicly indexed.
Imho its becoming increasingly clear that corporations will have to migrate to on prem / private cloud solutions in order to guarantee good data security practices.
via Alex
In the last year alone we had a series of data security blunders:
- xAI forgetting to add noindex on 400k conversations
- xAI uploading entire private repositories of thousands of clients to google cloud incl. .env files
- OpenAI forgetting to add noindex to shared conversations on 5k conversations
- Deepseek forgetting to add noindex to conversations (scope unknown)
- OpenAI unable to control its own sandbox environment, taking 3 days to notice that its own model was on the internet commiting felonies (the ChatGPT HuggingFace hack)
Alex further writes
I did a quick back of the envelope breakeven analysis.
Blackwell Nodes needed to run GLM5.2 at over 20tps output cost around 500k$US (if you can get one at all), and cost around 40$/h to rent from east asian private cloud providers. Electricity to run the node is around 300k$US in mid power markets. Hosted API services are available at lets say a blended rate of 2.5$/M tokens.
This means the onprem cut-even is at around 25k operating hours (ie. 96% utilization over 3 years).
The managed API breakeven is at 310 million tokens a day, which is impossible given the node can only produce 38 million tokens a day.
It is therefore not economically feasible (yet) to run on hardware. And its also clear that the frontier labs are still subsidizing the real costs of the models by over 80% (ie. Will have to quintuple prices to become profitable at current efficiencies and current energy markets)
Alex with another message linking to Anthropic's policy statement on open weight models
Cursor announces Cursor Start, a new ₹649/month (around USD 6.8/month) plan designed for Indian developers with local INR billing and UPI payments.
The tier offers generous daily access to Grok 4.5 and Composer models plus autonomous cloud agents, iOS app control, and workflow extensions for planning, building, and shipping code.
Positioned between the free plan and Pro, it targets India's large developer base to boost accessible agentic AI coding with market-specific pricing.
PS: Cursor needs to bring an Android version asap. India and other emerging markets have a much larger Android base
via Parth of Cryptique who asked
Hey folks, what tools are you guys using to monitor agents in production? Current state of finding agents issues and debugging is a mess in my opinion.
Fabien responded with
Braintrust for both traces and evals.
Also Datadog for correlated traces, but subpar UI as of right now (but they're iterating fast)
Naor responded with
Bifrost + signoz.
xAI Announces Grok 4.6 and 4.7 Releases Soon
xAI plans to release Grok 4.6 around August 7 with 1.5 trillion parameters and upgrades to supervised fine-tuning and reinforcement learning. Grok 4.7 follows weeks later as a 2.1 trillion-parameter model, outperforming 4.6 overall with better token efficiency despite slightly slower speed. This rapid pace follows Grok 4.5's launch on July 16, praised for coding and real-world tasks, and already aiding engineers at Tesla and SpaceX.
Grok Build + Cursor would IMHO be considered the Dynamic Duo of harnesses particular if Composer 3 also turns out to be stellar for its price and how much usage limits Grok/Composer get within Cursor particularly with their Starter and Pro plans
via Ben
Big News!
MCP is now stateless, making it easier to deploy and scale remote servers.
Now that MCP is stateless, you can deploy on serverless and edge infrastructure, or scale horizontally behind any load balancer.
Fish Audio announced a $52M seed round and public launch of S2.1 Pro TTS model, claiming voice cloning from 5 seconds of audio, 2x faster inference than Cartesia, 1/6th the cost of Eleven Labs, and word-level control over emotion, intonation, and pacing.
Their promotional video showcases low-latency real-time voice interactions via a simulated pizza order call, demonstrating natural expressiveness and responsiveness with the new model.
One year after starting as open-source, the company reached $21M ARR with enterprise clients including HeyGen, LiveKit, Retell, Sanas, and OpenArt running it in production, plus a free month of S2.1 Pro giveaway and a cost-cut guarantee for businesses.
Frontier company employees and execs have signed a petition urging the US to support an international effort to “deliberately pace” frontier AI development.
"We request that the U.S. government support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development"
Superwhisper announced a partnership with Cohere to integrate Transcribe, their 2B-parameter open-source speech recognition model, for fully offline local dictation in the voice-to-text app.
Transcribe delivers high accuracy with 5.35% average word error rate (1.25% on clean audio), supports 14 languages, outperforms Whisper Large v3, and handles custom vocabulary from the first use while keeping data private on-device.
AI Agent Harnesses Drive Huge Cost and Speed Differences
Composio evaluated Kimi K3 on 28 identical tasks across three agent harnesses—Claude Code, Hermes, and Kimi Code—revealing nearly identical success rates of 20-22/28 but dramatic differences in efficiency.
Token usage varied up to 30x, with medians of 61k for Kimi Code ($0.22 per task), 67k for Hermes ($0.28), and 340k for Claude Code ($2.00), driven mostly by input tokens at K3's pricing.
Hermes completed tasks fastest at 179 seconds median while Kimi Code led in token efficiency, underscoring that harness/framework design often drives agent costs and performance more than the model itself.
Consolidating news about various search and crawler API's typically used by agentic apps and harnesses
Firecrawl announced a major upgrade to its /search API, deploying a custom model that scores paragraphs, lists, and tables to return only the most relevant excerpts answering the query.
Agents using the new search achieve 94.7% accuracy on OpenAI's SimpleQA factuality benchmark—higher than competitors—while consuming 10x fewer tokens than full-page processing.
Octen web search API returns results in 62ms P50 with 6ms P50-P90 gap, several times faster than alternatives.
Costs $1 per 1,000 calls while leading on RTEB, DeepResearch Bench, FreshQA, and SimpleQA benchmarks.
Decomposes questions into dozens of sub-queries fired simultaneously for full source-backed reports in under 3 minutes.
Supports over 1M QPS with 5-minute indexing
Context.dev gives your AI agents and apps real-time access to structured web data, no brittle scraping infrastructure needed. Scrape any URL as clean markdown or HTML, extract brand data (logos, colors, fonts, socials) from any domain, crawl sitemaps, resolve transaction descriptors, and more.
PS: In a few of my harness I currently use Exa MCP + X API MCP. I'm testing out Octen in some harness's. Firecrawl is available as part of Nous Tool Gateway in Hermes so those who have a Nous Portal subscription automatically use Firecrawl for web_search/web_extract tool calls
and obviously I was living under a rock about TinyFish till today after I found out via SpaceXAI about them.
TinyFish doesn't charge for its search and fetch API 🤯, gives you 500 credits on signup so no brainer to give it a shot
OpenAI announces 80% price reduction for GPT-5.6 Luna and 20% for Terra, plus a fast mode for Sol delivering up to 2.5x speed at 2x standard price in the API.
Given all the previous posts where I highlighted a few programatic search and fetch providers I decided to create the below public Notion page that attempts to explain the benefits of paying for and using an external search/fetch provider(s) in an harness compared to the built-in search and fetch tools within an harness. In the article, I demonstrate by integrating the same search and fetch provider MCP's in different harness and using different models
I humbly appreciate any comments on this Notion page that I wrote
MiniMax launched H3, a unified multimodal generation model that processes text, images, video, and audio inputs together to output videos up to 15 seconds long at 2K resolution with native stereo sound.
H3 excels in complex instruction following, precise text/brand rendering, motion transfer, and controllable editing for commercial uses like advertising, e-commerce, and product design, while offering superior price-performance at under one-third the cost of mainstream 2K models.
The company plans to release H3 model weights openly soon to foster an open ecosystem, hardware compatibility, and custom versions, shifting from siloed specialized tasks toward general-purpose multimodal intelligence grounded in natural language.
DeepSeek announced the public beta launch of its DeepSeek-V4-Flash Official API, highlighting major upgrades to agent capabilities that now surpass the previous V4-Pro-Preview version on multiple benchmarks.
The accompanying benchmark table shows V4-Flash-0731 achieving large gains over its preview across agentic and coding tasks, including Terminal Bench 2.1 (82.7), DeepSWE (54.4), and DSBench-FullStack (68.7), putting it competitive with GLM-5.2 and Opus-4.8.
The update maintains the same model architecture and size as the preview, adds native support for Responses API format and Codex adaptation, while the full V4-Pro release is expected soon
Poolside released Desktop Assistant, a unified workspace for running and supervising coding agents across macOS, VS Code, and Visual Studio, addressing how long-horizon agents now operate for hours across repos rather than through simple chats.
Core features include vendor-agnostic support for agents like their Laguna models, Claude, Codex, or Gemini; parallel execution in isolated Git worktrees; and seamless handoffs that preserve full conversation context via the open Agent Client Protocol (ACP).
Developed from a year of internal daily use, the tool is available now for download, with the macOS version called Desktop Assistant and IDE versions as Poolside Assistant, enabling practical multi-agent workflows for complex software engineering tasks
ByteDance introduced Seedance 2.5, a video generation model focused on 30-second single-pass audio-video clips, supporting up to 30 images, 10 videos, and 10 audio references plus timestamp-level editing for precise control.
enterprise API access through BytePlus coming soon.
coming soon to Higgsfield.
Chamath Palihapitiya of Social Capital fame presents an AI stack diagram with physical scarcity at the bottom (land + power + shell, silicon, clouds, models) and customer proximity at the top (harnesses, applications), arguing margins will shift upward if models commoditize through context, workflows, and low switching costs.
He is heavily bullish on LPS (land/power/shell) for quick cash returns amid data center constraints and has acquired nearly 6GW of grid and behind-the-meter power capacity coming online through 2029 with partner Anita V. Lallian.
Chamath avoids silicon investments due to extreme complexity and capital waste, is cautious on clouds over KYC/alignment risks, questions model revenue durability from tokenmaxxing, and focuses his own efforts on harnesses via his $135M-funded 8090 platform for enterprise alpha and custom applications
Below is my personal website which aggregates links to many of my socials as well as the various content and community that I curate. Feel free to share this link to others who you think may find this content/community useful to them
The cover image of this newsletter via generated via the Seedream 5 Pro model within the Krea tool via the following prompt
Quaint rustic witch's cabin by the lake, autumn forest background, orange and honey colors, beautiful composition, magical, warm glowing lighting, cloudy, dreamy masterpiece, Nikon D610, photorealism, highly artistic, highly detailed, ultra high resolution, sharp focus, Mysterious
Over 800 subscribers