AI Digest Source Log
Append-only log of every source URL that already has an AIDigest/*.md article written about it. The scheduler checks this file first, before researching or writing anything, to avoid producing a duplicate article for a source already covered — faster and more reliable than re-scanning every article's frontmatter each run. AIDigest/*.md source: frontmatter fields remain the backstop/source of truth if this file and reality ever disagree.
Rolling 60-day window (added 2026-07-24): entries older than 60 days get pruned from this log each run — see ai-digest-scheduler step 8. This is a dedup index, not a permanent record; the articles themselves under AIDigest/YYYY/MM/DD/ are never affected by pruning here.
One line per article, newest at the bottom, format:
- [Title](slug.md) — <source-url> — YYYY-MM-DD
Log
-
Claude's Web Search Tool: Dynamic Filtering and Built-In Citations — https://platform.claude.com/docs/en/agents-and-tools/tool-use/web-search-tool.md — 2026-07-10
-
RFC 10008: HTTP Gets a Native QUERY Method — https://datatracker.ietf.org/doc/html/rfc10008 — 2026-07-10
-
MCP Goes Stateless: Beta SDKs for the 2026-07-28 Spec Revision — https://blog.modelcontextprotocol.io/posts/sdk-betas-2026-07-28/ — 2026-07-10
-
NVIDIA and LangChain Ship NemoClaw: An Open Stack for Deep Agents — https://blogs.nvidia.com/blog/nemotron-langchain-agents-open-stack/ — 2026-07-10
-
Singapore's MAS Publishes SAFR: Runtime Guardrails for Agentic Finance — https://www.mas.gov.sg/news/media-releases/2026/mas-partners-industry-to-develop-safeguards-for-ai-agents-in-finance — 2026-07-10
-
Prompt-to-Paper: An Agentic System That Writes (and Runs) Bioinformatics Research — https://arxiv.org/abs/2607.05456 — 2026-07-10
-
Memora: Microsoft Research's Harmonic Memory for Long-Horizon Agents — https://www.microsoft.com/en-us/research/blog/memora-a-harmonic-memory-representation-balancing-abstraction-and-specificity/ — 2026-07-10
-
Claude Reflect: Anthropic Ships a Usage-Habits Dashboard — https://techcrunch.com/2026/07/09/anthropics-new-claude-feature-is-quietly-selling-you-on-ai/ — 2026-07-10
-
US Military Health System Completes Global Rollout of Its Clinical AI Agent — https://www.medicaldaily.com/military-health-system-ai-ambient-listening-clinical-documentation-2026-476020 — 2026-07-10
-
MongoDB Ships Native Reranking and voyage-context-4 for In-Database RAG — https://www.mongodb.com/company/newsroom/press-releases/mongodb-delivers-accurate-ai-retrieval-wherever-enterprise-data-lives — 2026-07-10
-
LangChain Brings Recursive Language Models to Deep Agents to Fight Context Rot — https://www.langchain.com/blog/how-to-use-rlms-in-deep-agents — 2026-07-10
-
NVIDIA and Hugging Face Expand LeRobot with Isaac GR00T 1.7 and Isaac Teleop — https://blogs.nvidia.com/blog/hugging-face-lerobot-models-frameworks-open-robotics/ — 2026-07-10
-
Docker Makes the Case for microVM Isolation of Coding Agents — https://www.docker.com/blog/why-ai-agents-need-isolation/ — 2026-07-10
-
Elastic's Own Security Team Cuts Alert Triage to Under 3 Minutes with Agentic Workflows — https://www.elastic.co/security-labs/alert-triage-agentic-soc-elastic-workflows — 2026-07-10
-
Netflix's GenPage Replaces Multi-Stage Recommenders with One Generative Model — https://netflixtechblog.com/genpage-towards-end-to-end-generative-homepage-construction-at-netflix-77146fba8a08 — 2026-07-10
-
Simon Willison Finds Claude Sonnet 5's New Tokenizer Makes English ~1.4x More Expensive — https://simonwillison.net/2026/Jun/30/claude-sonnet-5/ — 2026-07-10
-
Safari Ships a Built-In MCP Server for Coding Agents to Inspect and Debug Web Pages — https://9to5mac.com/2026/07/01/safaris-new-mcp-server-lets-coding-agents-inspect-and-debug-websites/ — 2026-07-10
-
OpenAI Audit Finds Up to a Third of SWE-bench Verified Tasks Are Broken — https://openai.com/index/separating-signal-from-noise-coding-evaluations/ — 2026-07-10
-
EvoAgentBench Measures Whether Agents Can Transfer Learned Abilities Across Tasks — https://arxiv.org/abs/2607.05202 — 2026-07-10
-
AstraZeneca Researchers Cap Agent Memory at ~300 Tokens for Long-Horizon Drug R&D Modeling — https://arxiv.org/abs/2607.07666 — 2026-07-10
-
UpDoc Gets First FDA Clearance for a Patient-Facing Generative AI Medical Device — https://www.statnews.com/2026/07/02/fda-clearance-raises-questions-updoc-use-generative-ai-diabetes-treatment/ — 2026-07-10
-
Microsoft Commits $2.5B to Embed AI Engineers Directly Inside Enterprise Clients — https://techstartups.com/2026/07/02/microsoft-launches-frontier-company-with-2-5b-to-help-enterprises-choose-the-best-ai-models-and-maximize-roi/ — 2026-07-10
-
JPMorgan's AI Portfolio Agents Beat a 60/40 Benchmark in Two-Decade Backtest — https://www.bloomberg.com/news/articles/2026-07-09/jpmorgan-builds-ai-agents-that-beat-60-40-portfolio-in-backtests — 2026-07-10
-
CaixaBank and Visa Complete a Real-World AI-Agent-Initiated Payment — https://www.caixabank.com/en/headlines/news/caixabank-completes-its-first-transaction-initiated-by-an-artificial-intelligence-agent-in-collaboration-with-visa — 2026-07-10
-
Anthropic Finds a 'Global Workspace' Inside Claude's Reasoning — https://www.anthropic.com/research/global-workspace — 2026-07-11
-
Huawei's openJiuwen Ships AutoGenetic Memory, a Self-Evolving Store for AI Agents — https://www.techtimes.com/articles/319523/20260702/ai-agent-memory-learns-across-sessions-huawei-framework-ships-with-china-data-risk.htm — 2026-07-11
-
MemSyco-Bench Measures When Agent Memory Makes Models More Sycophantic — https://arxiv.org/abs/2607.01071 — 2026-07-11
-
EvoPolicyGym Benchmarks Agents on Iteratively Evolving Their Own Policies — https://arxiv.org/abs/2607.02440 — 2026-07-11
-
Mistral Confirms a New 'Fat but Sparse' Open-Weight Model Family — https://techcrunch.com/2026/07/04/what-is-mistral-ai-everything-to-know-about-the-openai-competitor/ — 2026-07-11
-
Kimi K2.7 Becomes the First Open-Weight Model in GitHub Copilot — https://github.blog/changelog/2026-07-01-kimi-k2-7-is-now-available-in-github-copilot/ — 2026-07-11
-
xAI Ships Grok 4.5, Its First Model Built for Coding and Agentic Work — https://x.ai/news/grok-4-5 — 2026-07-11
-
OpenAI Launches the GPT-5.6 Family: Sol, Terra, and Luna — https://openai.com/index/gpt-5-6/ — 2026-07-11
-
Google Ships ADK 2.0 for Go, Turning Agent Orchestration into an Explicit Graph — https://developers.googleblog.com/announcing-adk-go-20/ — 2026-07-11
-
Docker: Your Laptop Is Now a Production Environment — https://www.docker.com/blog/your-laptop-is-the-new-production-environment/ — 2026-07-11
-
Elastic Proposes a Four-Layer Governance Model for Autonomous Security Agents — https://www.elastic.co/blog/the-future-of-governing-ai-agents — 2026-07-11
-
Databricks Benchmarks Coding Agents on Its Own Multi-Million-Line Codebase — and Switches Its Default Model — https://www.databricks.com/blog/benchmarking-coding-agents-databricks-multi-million-line-codebase — 2026-07-11
-
GitHub Copilot's Agent Mode in VS Code Gets Native Browser Control at GA — https://github.blog/changelog/2026-07-01-browser-tools-for-github-copilot-in-vs-code-are-generally-available/ — 2026-07-11
-
UK's FCA Publishes the First Regulator-Led Review of AI in Retail Financial Services — https://www.fca.org.uk/news/press-releases/fca-publishes-landmark-review-impact-ai-retail-financial-services — 2026-07-11
-
HHS Signals a Federal Push Toward Autonomous Clinical AI Agents — https://distilinfo.com/2026/06/29/hhs-backs-autonomous-ai-agents-for-clinical-care/ — 2026-07-11
-
Takeda Signs Up to $600M AI Drug Discovery Deal with Insilico Medicine — https://www.eurekalert.org/news-releases/1134421 — 2026-07-11
-
Brook.ai Partners with SRHO to Bring Agentic Remote Care to 275+ Hospitals — https://www.prnewswire.com/news-releases/brookai-and-srho--the-national-association-partner-to-bring-ai-enabled-continuous-care-infrastructure-to-more-than-275-hospitals-nationwide-302818538.html — 2026-07-11
-
Baptist Health and Vega Health Launch AI Screening for Hospital-at-Home Eligibility — https://www.businesswire.com/news/home/20260709534039/en/Vega-Health-and-Baptist-Health-Bring-Actionable-AI-to-Hospital-at-Home-Set-Foundation-for-Long-Term-Collaboration — 2026-07-11
-
Cross River Bank Extends Stripe Partnership to Power Agentic Commerce Cards — https://www.businesswire.com/news/home/20260701428290/en/Cross-River-Expands-Stripe-Issuing-Partnership-to-Help-Power-Agentic-Commerce — 2026-07-11
-
Targa Telematics Targets Fleet Financiers with a New Agentic AI Platform — https://fintech.global/2026/07/07/targa-telematics-targets-fleet-financiers-with-agentic-ai/ — 2026-07-11
-
DeepFabric Reaches GA with 50+ AI Agents for Supply-Chain Operations — https://www.prnewswire.com/news-releases/deepfabric-announces-general-availability-of-its-ai-agent-platform-for-end-to-end-supply-chain-operations-302820223.html — 2026-07-11
-
Google Ships a Preview Agents API for Genkit, Its Open-Source Agent Framework — https://developers.googleblog.com/build-agentic-full-stack-apps-with-genkit/ — 2026-07-11
-
Claude Sonnet 5 Comes to Amazon Bedrock, and Bedrock Agents Enters Maintenance Mode — https://aws.amazon.com/blogs/machine-learning/introducing-claude-sonnet-5-on-aws-anthropics-most-capable-sonnet-model/ — 2026-07-11
-
MongoDB Moves Atlas Dedicated Clusters on AWS to Gen2 ARM Infrastructure — https://www.mongodb.com/company/blog/product-release-announcements/introducing-atlas-gen2-on-aws-m30-dedicated-clusters — 2026-07-11
-
LangChain Releases OpenWiki, an Agent That Writes and Maintains Docs for Coding Agents — https://www.langchain.com/blog/introducing-openwiki-an-open-source-agent-for-repo-documentation — 2026-07-11
-
NVIDIA Ties Its Economics to Revenue-Sharing Deals to Unblock the AI Compute Bottleneck — https://blogs.nvidia.com/blog/nvidia-unlocks-ai-compute-at-scale-capital-partners-to-power-ai-infrastructure-buildout/ — 2026-07-11
-
Oracle Adds Custom Extraction and Hybrid Search to Its AI Agent Memory Core — https://blogs.oracle.com/developers/whats-new-in-oracle-ai-agent-memory-custom-extraction-hybrid-search-and-more-control — 2026-07-11
-
OpenAI Launches GPT-Live, a Full-Duplex Voice Model Family for ChatGPT — https://openai.com/index/introducing-gpt-live/ — 2026-07-11
-
ToolFailBench Separates Why LLM Agents Fail at Tool Use, Not Just Whether They Do — https://arxiv.org/abs/2607.04686 — 2026-07-11
-
Study Finds Mixed-Personality AI Agent Teams Outperform Uniform Ones on Coding Tasks — https://arxiv.org/abs/2607.05659 — 2026-07-11
-
Study Finds MCP, A2A, and ACP All Lack Native Multi-Party Governance Controls — https://arxiv.org/abs/2606.31498 — 2026-07-11
-
GitHub Copilot Becomes a Native Agent Option Inside JetBrains AI Assistant — https://github.blog/changelog/2026-06-30-copilot-agent-is-now-available-in-jetbrains-ai-assistant/ — 2026-07-11
-
GitHub Lets Enterprises Stream Copilot Agent Sessions to Their Own SIEM — https://github.blog/changelog/2026-07-02-copilot-agent-session-streaming-is-now-in-public-preview/ — 2026-07-11
-
Elastic Finds Signs of LLM-Assisted Development in a Mexican Banking Fraud Toolkit — https://www.elastic.co/security-labs/mexican-banking-fraud-scmbanker-ref6045 — 2026-07-11
-
AWS CloudFormation Express Mode Cuts Deploy Feedback Loops for AI-Agent-Driven Infra Work — https://aws.amazon.com/blogs/aws/accelerate-your-infrastructure-deployments-by-up-to-4x-with-aws-cloudformation-express-mode/ — 2026-07-11
-
Philips Launches Alturion, an AI-Guided Ultrasound System for High-Volume Clinics — https://www.philips.com/a-w/about/news/archive/standard/news/press/2026/philips-introduces-alturion-ultrasound-system-with-ai-powered-workflows-for-high-volume-clinical-environments.html — 2026-07-11
-
Abrigo Launches an Agentic AI Platform to Automate Bank Lending Workflows — https://www.businesswire.com/news/home/20260708110699/en/Abrigo-Launches-Agentic-AI-Platform — 2026-07-11
-
hyperexponential Launches an Agentic Underwriting Workbench for Commercial P&C Insurance — https://www.reinsurancene.ws/hyperexponential-launches-hyperoperator-to-automate-commercial-pc-underwriting-workflows/ — 2026-07-11
-
Rackspace and Palantir Launch a Joint Framework for Enterprise AI in Regulated Industries — https://www.globenewswire.com/news-release/2026/07/09/3324714/0/en/Rackspace-Technology-Launches-Operating-Framework-with-Palantir-for-Regulated-Enterprises-to-Accelerate-Enterprise-AI-in-Production.html — 2026-07-11
-
ProGlove's Lessons from Scaling a Serverless SaaS Platform to 1 Million Lambda Functions — https://aws.amazon.com/blogs/architecture/lessons-learned-from-scaling-to-1-million-lambda-functions/ — 2026-07-11
-
'Friendly Fire': Prompt Injection Turns Claude Code and Codex Into Their Own Malware Vector — https://ainowinstitute.org/publications/friendly-fire-exploit-brief — 2026-07-11
-
OpenAI Merges Codex Into ChatGPT Desktop and Launches an Autonomous 'ChatGPT Work' Agent — https://openai.com/index/chatgpt-for-your-most-ambitious-work/ — 2026-07-11
-
Study Finds Runtime-Based Coding-Agent Benchmarks Are Full of Measurement Noise — https://arxiv.org/abs/2607.01211 — 2026-07-11
-
Study Formalizes the Bias-vs-Noise Tradeoff in LLM-as-Judge Evaluation Setups — https://arxiv.org/abs/2607.00304 — 2026-07-11
-
Databricks Brings OpenAI's Codex Natively Onto Its Platform, Governed by Unity AI Gateway — https://www.databricks.com/blog/openai-and-databricks-dais-2026-making-enterprise-ai-real — 2026-07-11
-
A Developer's Take on Anthropic's 'When AI Builds Itself': Execution Is Automatable, Judgment Isn't — https://dev.to/hemapriya_kanagala/reading-anthropics-when-ai-builds-itself-changed-how-i-think-about-ai-and-software-engineering-3eh — 2026-07-11
-
Meta Ships Muse Spark 1.1, Its First Paid Agentic Model — https://www.marktechpost.com/2026/07/09/meta-superintelligence-labs-releases-muse-spark-1-1/ — 2026-07-12
-
LangChain Ships OpenWiki Brains, Proactive Wiki-Style Memory for Agents — https://www.langchain.com/blog/introducing-openwiki-brains-general-purpose-wiki-memory-for-agents — 2026-07-12
-
Meta AI's 'Remember When It Matters' Gives Agents an Active Memory Manager — https://arxiv.org/abs/2607.08716 — 2026-07-12
-
A New Severity Scale Moves Agent Red-Teaming Past Binary Attack Success — https://arxiv.org/abs/2607.07474 — 2026-07-12
-
MUTE Teaches Multi-Agent Systems to Unlearn Wasteful Communication — https://arxiv.org/abs/2607.03473 — 2026-07-12
-
PiSAs Benchmarks Whether Multi-User Agents Leak Information Across Users — https://arxiv.org/abs/2607.05318 — 2026-07-12
-
What Do Coding Agents Actually Change? A Taxonomy of 1,254 PR Diffs — https://arxiv.org/abs/2607.05666 — 2026-07-12
-
Simon Willison: 'Understanding Is the New Bottleneck' for Agentic Coding — https://simonwillison.net/2026/Jul/2/understand-to-participate/ — 2026-07-12
-
Simon Willison: Let the Agent Decide When to Delegate to a Cheaper Subagent — https://simonwillison.net/2026/jul/3/judgement/ — 2026-07-12
-
Can LLMs Design Valid Network Topologies from Plain-Language Intent? — https://arxiv.org/abs/2607.00292 — 2026-07-12
-
OmniPilot Predicts LLM Serving Cost Across Heterogeneous GPU Clusters — and Knows When to Abstain — https://arxiv.org/abs/2607.01579 — 2026-07-12
-
Abu Dhabi and MIT's Koch Institute Team Up on AI-Driven Cancer Research — https://www.prnewswire.com/news-releases/doh-and-mit-koch-institute-partner-to-advance-ai-driven-cancer-research-302817573.html — 2026-07-12
-
Telepatía AI Raises $42M to Scale Its AI Clinical Assistant Across Latin America — https://colombiaone.com/2026/07/09/telepatia-ai-42-million-funding/ — 2026-07-12
-
Aily Labs and AWS Bring Agentic Decision Intelligence to the Fortune 500 — https://www.prnewswire.com/news-releases/aily-labs-and-aws-announce-strategic-partnership-to-accelerate-ai-decision-intelligence-across-the-fortune-500-302817247.html — 2026-07-12
-
Ben Bernanke Joins Anthropic's Long-Term Benefit Trust — https://www.bloomberg.com/news/articles/2026-07-09/former-fed-chairman-ben-bernanke-joins-anthropic-oversight-trust — 2026-07-12
-
PNC Rebuilds Its Mobile Banking App Around Embedded Generative AI — https://www.prnewswire.com/news-releases/pncs-new-mobile-app-delivers-an-intuitive-personalized-experience-to-help-clients-seamlessly-manage-their-day-to-day-financial-lives-302819063.html — 2026-07-12
-
Regnology Acquires Fed Reporter, Extending Agentic Regulatory Reporting to 4,000+ US Banks — https://www.businesswire.com/news/home/20260701405645/en/Regnology-to-Acquire-Fed-Reporter-Accelerating-U.S.-Leadership-and-Advancing-Regulatory-Modernization — 2026-07-12
-
Kord Raises £6.4M to Fight AI-Driven Fraud in Regulated Industries — https://fintech.global/2026/07/09/kord-raises-6-4m-to-fight-ai-fraud-in-regulated-sectors/ — 2026-07-12
-
GIM Raises $20M as Agentic AI Investing Moves From Research to Live Execution — https://fintech.global/2026/07/10/gim-raises-20m-to-scale-agentic-ai-investing/ — 2026-07-12
-
Databricks Proposes 'Decision Execution Platforms' to Close the Loop From Dashboard to Action — https://www.databricks.com/blog/beyond-dashboards-introducing-decision-execution-platforms — 2026-07-12
-
AWS's Field Guide to MCP Tool Design: Granularity, Schemas, and Tradeoffs — https://aws.amazon.com/blogs/machine-learning/mcp-tool-design-practical-approaches-and-tradeoffs/ — 2026-07-12
-
AWS Launches Claude Apps Gateway for Centralized Coding-Agent Governance — https://aws.amazon.com/blogs/machine-learning/introducing-claude-apps-gateway-for-aws/ — 2026-07-12
-
GitHub Models, GitHub's Free AI Model Playground, Is Being Fully Retired — https://github.blog/changelog/2026-07-01-github-models-is-being-fully-retired-on-july-30-2026/ — 2026-07-12
-
MongoDB Brings Search and Vector Search GA to Self-Managed Deployments — https://www.mongodb.com/company/blog/product-release-announcements/supercharge-self-managed-apps-search-vector-search-capabilities — 2026-07-12
-
How Box AI Went Agent-Native With LangChain Deep Agents — https://www.langchain.com/blog/building-box-ai-how-an-enterprise-content-platform-went-ai-native-with-deep-agents — 2026-07-12
-
NVIDIA's Vera CPU Bets on Single-Threaded Performance for the Agent Loop — https://blogs.nvidia.com/blog/nvidia-vera-max-single-threaded-cpu-at-scale/ — 2026-07-12
-
Google Shows Elastic Training Recovering From a Mid-Training TPU Failure in Seconds — https://developers.googleblog.com/we-terminated-a-tpu-mid-training-and-it-recovered-in-seconds-introduction-to-elastic-training-with-maxtext/ — 2026-07-12
-
LiteRT.js Brings Google's On-Device Inference Engine to the Browser — https://developers.googleblog.com/litertjs-googles-high-performance-web-ai-inference/ — 2026-07-12
-
NVIDIA's Open Models Become ICML 2026's Research Infrastructure — https://blogs.nvidia.com/blog/open-models-icml-2026/ — 2026-07-13
-
AWS Publishes a Serverless A2A Gateway for Agent Discovery and Access Control — https://aws.amazon.com/blogs/machine-learning/building-a-serverless-a2a-gateway-for-agent-discovery-routing-and-access-control/ — 2026-07-13
-
AWS and Mistral Demo a Production Ecommerce MCP Server on Bedrock AgentCore — https://aws.amazon.com/blogs/machine-learning/building-and-connecting-a-production-ready-ecommerce-mcp-server-using-amazon-bedrock-agentcore-and-mistral-ai-studio/ — 2026-07-13
-
Databricks Ships Feature Views to Unify Training and Serving Pipelines — https://www.databricks.com/blog/introducing-feature-views — 2026-07-13
-
GitHub Copilot Adds OpenAI's GPT-5.6 Sol, Terra, and Luna to Its Model Picker — https://github.blog/changelog/2026-07-09-openais-gpt-5-6-sol-terra-and-luna-are-now-available-in-github-copilot/ — 2026-07-13
-
VS Code Copilot's June 2026 Update Adds Cost Visibility and a More Autonomous Autopilot — https://github.blog/changelog/2026-07-08-github-copilot-in-visual-studio-code-june-2026-releases/ — 2026-07-13
-
NHS England Details £10B AI Rollout: National Triage Tool and Ambient Notetaking — https://www.england.nhs.uk/2026/07/nhs-accelerates-artificial-intelligence-rollout-to-cut-waiting-times-and-improve-care-for-millions/ — 2026-07-13
-
UN's First Global AI Science Panel Warns on Sycophantic Chatbots and Mental Health Harms — https://www.un.org/independent-international-scientific-panel-ai/en/preliminary-report — 2026-07-13
-
Pearl Health Raises $110M to Scale AI for Value-Based Medicare Care — https://www.prnewswire.com/news-releases/pearl-health-raises-110-million-to-expand-its-ai-platform-helping-providers-deliver-better-outcomes-at-lower-cost-for-medicare-patients-302820795.html — 2026-07-13
-
Geisinger Study Links AI-Guided Cancer-Screening Outreach to Lower Mortality — https://www.eurekalert.org/news-releases/1135090 — 2026-07-13
-
Stanford's CANVAS Reads Routine Pathology Slides to Predict Immunotherapy Resistance — https://news.stanford.edu/stories/2026/07/ai-platform-canvas-tumor-pathology — 2026-07-13
-
Terence Tao Revives a 27-Year-Old Java Applet With a Coding Agent, in Hours — https://terrytao.wordpress.com/2026/07/11/old-and-new-apps-via-modern-coding-agents/ — 2026-07-13
-
JetBrains Launches a Kotlin-Specific Benchmark for AI Coding Agents — https://blog.jetbrains.com/kotlin/2026/07/introducing-the-kotlin-benchmark-evaluate-ai-coding-agents-on-real-world-kotlin-tasks/ — 2026-07-13
-
DeepInfra Opens Its First International Data Center, in Toronto — https://www.globenewswire.com/news-release/2026/07/08/3324125/0/en/DeepInfra-Expands-AI-Inference-Capacity-with-First-International-Data-Center-in-Toronto.html — 2026-07-13
-
Together AI Raises $800M at $8.3B Valuation as Open-Weight Inference Demand Triples — https://www.businesswire.com/news/home/20260701243402/en/Together-AI-Raises-$800-Million-at-$8.3-Billion-Valuation-to-Make-Frontier-AI-Accessible-to-All — 2026-07-13
-
Nathan Lambert: US Regulation, Not Capability, Is Now the Ceiling on Open Models — https://www.interconnects.ai/p/6-months-to-live-for-open-models — 2026-07-13
-
Memory Architecture, Not Channel Capacity, Drives Emergent Language in LLM Agents — https://arxiv.org/abs/2607.00233 — 2026-07-13
-
PACE: Predicting Expensive Agentic Benchmark Scores From Cheap Proxy Evals — https://arxiv.org/abs/2607.02032 — 2026-07-13
-
SovereignPA-Bench Tests Whether Personal AI Agents Respect User Consent Under Pressure — https://arxiv.org/abs/2607.05363 — 2026-07-13
-
Single-Rollout Sampling Stabilizes Asynchronous RL for Long-Horizon Agents — https://arxiv.org/abs/2607.07508 — 2026-07-13
-
From Prompts to Contracts: A Harness-Engineering Pattern for Auditable Enterprise Agents — https://arxiv.org/abs/2607.08028 — 2026-07-13
-
TRACE Watermarks Agent Trajectories to Survive Tampering by the Party That Logs Them — https://arxiv.org/abs/2607.08400 — 2026-07-13
-
IdeaGene-Bench Tests Whether LLMs Can Reason About How Scientific Ideas Inherit and Recombine — https://arxiv.org/abs/2607.08758 — 2026-07-13
-
Cisco Rolls Out AI Agents to All 90,000 Employees, Builds a 'CFO Cockpit' — https://fortune.com/2026/07/01/cisco-cfo-ai-agents-finance-employees-mark-patterson/ — 2026-07-13
-
Bank of England Floats a Market-Wide 'Kill Switch' for Agentic Trading — https://www.techtimes.com/articles/319549/20260702/half-finance-firms-run-autonomous-ai-traders-bank-england-proposes-market-kill-switch.htm — 2026-07-13
-
ECB and ESRB Warn Banks on Systemic Cyber Risk From Frontier AI Models — https://www.regulationtomorrow.com/2026/07/esrb-warning-on-frontier-ai-models-and-ecb-writes-to-significant-institutions/ — 2026-07-13
-
Ex-DeepMind Poker AI Team Raises Series A for AI Trading Agents at $500M+ Valuation — https://siliconangle.com/2026/07/01/ai-stock-trading-startup-equilibre-raises-funding-500m-valuation/ — 2026-07-13
-
Webull Launches Vega Portfolio, an AI Portfolio Advisor for Retail Investors — https://www.prnewswire.com/news-releases/webull-launches-ai-powered-portfolio-advisor-at-new-york-event-302815800.html — 2026-07-13
-
NY Fed Uses LLMs to Build the Largest-Ever Database of US Bank Runs — https://libertystreeteconomics.newyorkfed.org/2026/07/using-ai-to-let-history-speak-about-bank-runs/ — 2026-07-13
-
Advisor360° Launches a 'Zero-Click' AI Agent That Preps Advisor Client Meetings — https://www.advisor360.com/advisor360-launches-meeting-prep-an-ai-agent-for-financial-advisors — 2026-07-13
-
Sixfold Survey: 99% of Underwriting Executives More Optimistic About AI, but Talent Anxiety Persists — https://fintech.global/2026/07/08/86-of-underwriters-more-excited-about-ai-future-sixfold/ — 2026-07-13
-
Skyro Takes AI-Underwritten Digital Credit Nationwide in the Philippines — https://www.manilatimes.net/2026/07/13/tmt-newswire/media-outreach-newswire/skyro-rolls-out-reusable-digital-credit-across-the-philippines-explores-opportunities-in-southeast-asian-markets/2383017 — 2026-07-13
-
Candidly Adds AI Guidance for New 'Invest America' Tax-Advantaged Child Accounts — https://www.businesswire.com/news/home/20260702063358/en/Candidly-Expands-Its-AI-Financial-Guidance-Platform-with-Six-New-Capabilities-Including-Guidance-for-Trump-Accounts — 2026-07-13
-
Vestmark Opens a Dedicated AI Research and Development Center in Boston — https://www.planadviser.com/ai-product-service-launches-7-6-2026/ — 2026-07-13
-
Kitces: Salesforce, RightCapital, and YCharts Ship AI Features as Incumbents 'Strike Back' — https://www.kitces.com/blog/the-latest-in-financial-advisortech-july-2026-salesforce-rightcapital-ycharts-ai-news/ — 2026-07-13
-
Revolut X Opens Crypto Trading to Third-Party AI Assistants, With Human Approval Required — https://en.cryptonomist.ch/2026/07/10/revolut-ai-crypto-trading/ — 2026-07-13
-
Former Mayo Clinic AI Governance Director Alleges a 67% Clinical AI Error Rate Was Covered Up — https://medcitynews.com/2026/07/mayo-clinic-ai-lawsuit/ — 2026-07-14
-
Palantir Lands Its First Publicly Named Latin American Customer: Mexico's Largest Insurer — https://www.businesswire.com/news/home/20260707514163/en/Palantir-Expands-Its-Presence-in-Mexico-and-Strengthens-Its-AI-Offering-in-the-Insurance-Sector-with-GNP-Seguros — 2026-07-14
-
Companies That Cut Staff for AI Are Quietly Rehiring — Ford Says the Mistake Cost Hundreds of Millions — https://www.cnbc.com/2026/07/01/employers-who-laid-off-workers-for-ai-are-reversing-their-decisions.html — 2026-07-14
-
Anthropic Finds Claude's Values Shift by Model Version and Input Language — https://www.anthropic.com/research/claude-values-models-languages — 2026-07-14
-
Google Migrates Its Android Coding Leaderboard to the Open Harbor Framework — Claude Fable 5 Leads at 84.5% — https://android-developers.googleblog.com/2026/07/android-bench-llm-measurement.html — 2026-07-14
-
AWS Open-Sources Loom, a Reference Blueprint for Governed Enterprise Agents — https://aws.amazon.com/blogs/opensource/building-secure-ai-agents-at-scale-introducing-loom-for-aws/ — 2026-07-14
-
Anyshift's AI Agent Gets Read-Only Elasticsearch Access to Speed Up Incident Response — https://www.elastic.co/blog/anyshift-partnership — 2026-07-14
-
AgentKGV Cuts RAG Search Calls in Half While Improving Knowledge-Graph Fact Verification — https://arxiv.org/abs/2607.09092 — 2026-07-14
-
Prismata Confines Cross-Site Prompt Injection in Web Agents Without Developer Annotations — https://arxiv.org/abs/2607.08147 — 2026-07-14
-
Long-Horizon-Terminal-Bench Grades CLI Agents on Progress, Not Just Pass/Fail — https://arxiv.org/abs/2607.08964 — 2026-07-14
-
A Game-Theoretic Multi-Agent Framework Cuts Hallucination by Turning Training Into a Team Game — https://arxiv.org/abs/2607.08403 — 2026-07-14
-
PERFOPT-Bench Tests Whether Coding Agents Can Find Real Performance Bugs, Not Just Measurement Noise — https://arxiv.org/abs/2607.07744 — 2026-07-14
-
An MCP-Orchestrated Agent Pipeline Turns Legacy Compliance Docs Into Machine-Readable OSCAL — https://arxiv.org/abs/2607.08288 — 2026-07-14
-
What Actually Makes a Bug Report Useful to an AI Repair Agent? — https://arxiv.org/abs/2607.09553 — 2026-07-14
-
Do Deep-Research Pipelines Need a Frontier Model Just to Verify Citations? — https://arxiv.org/abs/2607.08700 — 2026-07-14
-
'Failure as a Process': A Large-Scale Anatomy of How CLI Coding Agents Actually Fail — https://arxiv.org/abs/2607.09510 — 2026-07-14
-
Agora Turns Agent Task Allocation Into an Auction So Work Goes to the Most Competent Model, Not the Most Confident One — https://arxiv.org/abs/2607.09600 — 2026-07-14
-
Two Layers of the Same AWS Agentic Stack: Workshop Blueprint vs. Loom Governance Platform — https://github.com/awslabs/loom — 2026-07-15
-
Docker: The Runtime Is Where Agent Trust Is Won — https://www.docker.com/blog/ai-engineer-worlds-fair-2026-the-runtime-is-where-agent-trust-is-won/ — 2026-07-15
-
GitHub Copilot in Visual Studio Adds an MCP Trust Layer and Ships C++ Modernization at GA — https://github.blog/changelog/2026-07-14-github-copilot-in-visual-studio-june-update/ — 2026-07-15
-
Google Ships Gemma 4 E2B, Tuned to Run Natively on the Pixel 10's TPU — https://developers.googleblog.com/unlocking-the-next-era-of-on-device-ai-with-google-tensor-and-pixel/ — 2026-07-15
-
Four Coding Agents, One Scaffold-to-PR Task: Mistral Vibe, Claude Code, Cursor, and Codex Scored Head-to-Head — https://www.marktechpost.com/2026/07/14/mistral-vibe-for-code-vs-claude-code-vs-cursor-vs-codex-four-agents-scored-on-one-scaffold-to-pr-task/ — 2026-07-15
-
PrismML Compresses a 27B Model Onto a Phone With Ternary and 1-Bit Builds of Qwen3.6 — https://www.marktechpost.com/2026/07/14/prismml-releases-bonsai-27b-1-bit-and-ternary-builds-of-qwen3-6-27b-that-run-on-laptops-and-phones/ — 2026-07-15
-
Concho AI Turns Enterprise Codebases Into an MCP-Exposed Knowledge Layer for Coding Agents — https://siliconangle.com/2026/07/14/concho-ai-turns-enterprise-codebases-knowledge-layer-ai-agents/ — 2026-07-15
-
Boundless Repurposes 4,000 Crypto-Proving GPUs Into AI Inference Capacity — https://siliconangle.com/2026/07/14/boundless-taps-idle-crypto-gpus-cut-ai-inference-costs/ — 2026-07-15
-
Demis Hassabis Calls for a FINRA-Style Independent Standards Body to Test Frontier AI Models — https://demishassabis.substack.com/p/a-framework-for-frontier-ai-and-the-dawning-of-a-new-age — 2026-07-15
-
JPMorgan's CFO: Stop Using Expensive Frontier Models for Easy Tasks — https://www.bloomberg.com/news/articles/2026-07-14/jpmorgan-urges-staff-to-avoid-deploying-pricey-ai-for-easy-tasks — 2026-07-15
-
StructAgent Doubles Long-Horizon Computer-Use Success by Replacing Raw History With Causal State — https://arxiv.org/abs/2607.11388 — 2026-07-15
-
A Training Recipe Extends LLM Context With Constant-Memory Associative Recurrence, Not More Attention — https://arxiv.org/abs/2607.11614 — 2026-07-15
-
LLMs Spontaneously Reproduce Human Irrationality in Route-Choice Decisions — https://arxiv.org/abs/2607.11632 — 2026-07-15
-
Databricks Becomes a Day-Zero Home for Thinking Machines Lab's First Open-Weights Model — https://www.databricks.com/blog/inkling-thinking-machines-lab-now-databricks — 2026-07-16
-
AWS Ships a Reference Build for a Zero-Orchestration Image-Editing Agent on Bedrock AgentCore — https://aws.amazon.com/blogs/machine-learning/build-a-serverless-image-editing-agent-with-amazon-bedrock-agentcore-harness/ — 2026-07-16
-
NVIDIA Shrinks Jetson Thor Into T3000 and T2000, Chasing Mainstream Robotics Price Points — https://blogs.nvidia.com/blog/jetson-thor-robotics-edge-ai-agent/ — 2026-07-16
-
GitHub Gives Enterprises a Cost Preview Before Code Quality Becomes a Paid Product — https://github.blog/changelog/2026-07-13-github-code-quality-license-estimate-in-public-preview/ — 2026-07-16
-
NVIDIA and Noetra Plan a 27,500-GPU Vera Rubin AI Factory to Anchor Japan's National AI Push — https://blogs.nvidia.com/blog/japan-ecosystem-2026/ — 2026-07-16
-
Weill Cornell's EmulatRx Puts a Five-Agent AI Team in Charge of Designing Clinical Trials — https://www.nature.com/articles/s41467-026-74501-2 — 2026-07-16
-
Mayo Clinic Now Runs 150 AI Models — Including One That Cuts Chart Review by Up to 30 Minutes — https://www.cnn.com/2026/07/16/tech/mayo-clinic-ai-healthcare — 2026-07-16
-
Insilico's Second Deal With China Medical System in Three Months Puts a Number on AI Drug Discovery Speed — http://www.prnewswire.com/news-releases/deepening-collaboration-in-ai-powered-rd-acceleration-insilico-medicine-and-cms-announce-additional-collaborations-in-cns-diseases-302823160.html — 2026-07-16
-
Survey: 74% of C-Suite Leaders Admit They Overstated Confidence in Their Own AI Strategy — https://www.globenewswire.com/news-release/2026/07/15/3327864/0/en/C-Suite-Leaders-Admit-Overstating-Confidence-in-Their-AI-Strategy-New-Survey-Finds.html — 2026-07-16
-
Starling Bank Cuts 130 Jobs, Says AI Adoption Is Part of Why — https://www.pymnts.com/news/banking/2026/starling-bank-cuts-130-jobs-amid-ai-adoption-and-restructuring/ — 2026-07-16
-
Float Raises €4.5M to Turn a Revenue-Financing Lender Into an AI-Native Finance Platform — https://fintech.global/2026/07/15/float-lands-e4-5m-series-a-to-close-europes-tech-funding-gap/ — 2026-07-16
-
Glia and Alloy Labs Release a Free Playbook to Fix Why 80% of Banks' AI Spend Isn't Paying Off — https://www.finopotamus.com/post/glia-and-alloy-labs-launch-interactive-kit-to-help-banking-leaders-plan-for-2027 — 2026-07-16
-
A Cognitive-Structured Multimodal Agent Replaces Raw Visual History With Episodic Memory — https://arxiv.org/abs/2607.08497 — 2026-07-16
-
xAI Open-Sources Grok Build While the Code That Exfiltrated Repos Stays In — https://www.techtimes.com/articles/320671/20260716/grok-build-open-sourced-after-covert-upload-code-exfiltrate-repos-stays.htm — 2026-07-16
-
What LLM Agents Say When No One Is Watching: Debate Study Finds Systematic Public/Private Divergence — https://arxiv.org/abs/2607.02507 — 2026-07-16
-
AlphaEvolve Reaches General Availability, With BASF, Klarna, and JetBrains as Proof Points — https://cloud.google.com/blog/products/ai-machine-learning/alphaevolve-is-available-for-everyone — 2026-07-16
-
PolyWorkBench Finds Long-Horizon LLM Agents Buckle Under Multilingual Workflows — https://arxiv.org/abs/2607.06008 — 2026-07-16
-
Agentic IoT: A Survey Maps the Path From Sensing Infrastructure to Cognitive Agent Ecosystems — https://arxiv.org/abs/2607.04219 — 2026-07-16
-
Verifiable Literate Programming Gives Non-Expert Coders a Way to Actually Review AI-Generated Code — https://arxiv.org/abs/2607.02333 — 2026-07-16
-
Structured Sparse Autoencoders Fix Fragmented Concepts in Vision-Language Models — https://arxiv.org/abs/2607.08605 — 2026-07-16
-
Moonshot AI Ships Kimi K3, a 2.8-Trillion-Parameter Open MoE Model With 1M-Token Context — https://www.marktechpost.com/2026/07/16/moonshot-ai-releases-kimi-k3-a-2-8-trillion-parameter-open-moe-model-with-kimi-delta-attention-and-1m-context/ — 2026-07-17
-
OpenAI's GPT-Red Trains Itself to Attack GPT-5.6, Beating Human Red-Teamers 84% to 13% — https://openai.com/index/unlocking-self-improvement-gpt-red/ — 2026-07-17
-
Terence Tao Builds Two More Math Visualization Apps With Coding Agents — https://terrytao.wordpress.com/2026/07/16/two-more-apps-visualizing-the-zeta-process-and-the-motions-of-the-heavens/ — 2026-07-17
-
Google Proposes Treating Agent Prompts Like Compiled Build Artifacts, Not Hand-Edited Text — https://developers.googleblog.com/building-scalable-ai-agents-with-modular-prompt-transpilation/ — 2026-07-17
-
Docker's Case for Why AI Agent Safety Is an Infrastructure Problem, Not a Model Problem — https://www.docker.com/blog/what-are-ai-agents/ — 2026-07-17
-
Databricks Ships Apache Spark 4.2 With a Native Semantic Layer and Built-In Geospatial Types — https://www.databricks.com/blog/introducing-apache-spark-42 — 2026-07-17
-
DeepSWE Builds a Coding-Agent Benchmark Designed to Resist Pretraining Contamination — https://arxiv.org/abs/2607.07946 — 2026-07-17
-
Lilian Weng: Recursive AI Self-Improvement Starts With the Harness, Not the Weights — https://lilianweng.github.io/posts/2026-07-04-harness/ — 2026-07-17
-
GitHub Admits Its Copilot Code Review Got Worse Before It Got Better — Here's the Fix — https://github.blog/ai-and-ml/github-copilot/better-tools-made-copilot-code-review-worse-heres-how-we-actually-improved-it/ — 2026-07-17
-
A New Paper Argues Agent Memory Should Live Inside the Reasoning Loop, Not Behind a Network Call — https://arxiv.org/abs/2607.05690 — 2026-07-17
-
NapMem Turns Agent Memory Into Something the Model Actively Navigates, Not Passively Receives — https://arxiv.org/abs/2607.05794 — 2026-07-17
-
Kleiner Perkins and John Doerr Bet on Embedding Startups Inside UCSF Health's Actual Workflows — https://fortune.com/2026/07/15/kleiner-perkins-john-doerr-and-ucsf-health-investment-healthcare-ai-mayo-clinic-venture-capital/ — 2026-07-17
-
SEC Names 'AI-Washing' a Top FY2026 Exam Priority as Only 24% of Firms Have a Vendor Policy — https://fintech.global/2026/07/15/sec-turns-up-heat-on-ai-and-retail-push-in-private-funds/ — 2026-07-17
-
Insilico Medicine Extends Its AI Drug Pipeline Into Manufacturing With a Potential $2.5B Bora Deal — https://www.techtimes.com/articles/320594/20260715/generative-ai-moves-from-drug-discovery-to-drug-factories-25b-bora-deal.htm — 2026-07-17
-
Visa Will Let Banks Embed a White-Label AI Financial Assistant Directly Into Their Apps — https://www.pymnts.com/news/artificial-intelligence/2026/visa-readies-rollout-ai-financial-assistant-banking-apps/ — 2026-07-17
-
WHO Europe: Two-Thirds of Countries Deploy Hospital AI, but Only 8% Have a Governance Strategy — https://www.euronews.com/health/2026/07/16/europe-needs-to-catch-up-with-the-ai-surge-in-hospitals-who-says — 2026-07-17
-
"Loop Engineering" Emerges as a Named Discipline: Two Independent Takes From Addy Osmani and LangChain — https://addyosmani.com/blog/loop-engineering/ — 2026-07-19
-
Claude Code Shipped a Silent Auto-Continue Timer — Anthropic Admits It Missed the Bar — https://www.olafalders.com/2026/07/17/claude-code-anatomy-of-a-misfeature/ — 2026-07-19
-
Terminal-Bench 2.1: Kimi K3 Lands #2 Overall, First Among Open-Weight Models — https://llm-stats.com/benchmarks/terminal-bench-2.1 — 2026-07-19
-
Kimi K3's Docs Confirm Three Reasoning-Effort Tiers, Mapped to Task Type — https://www.techtimes.com/articles/320937/20260718/kimi-k3-adds-standard-high-reasoning-modes-documentation-maps-three-effort-tiers.htm — 2026-07-19
-
OpenAI Codex Now Encrypts Sub-Agent Instructions — Locking Developers Out of Their Own Audit Trail — https://www.theregister.com/ai-and-ml/2026/07/15/openai-hides-codex-agent-instructions-behind-encryption-leaving-developers-in-the-dark/5271484 — 2026-07-19
-
Apple's Selective Persistent Memory Cuts Agent Task Time 14x by Forgetting Reasoning, Keeping Facts — https://arxiv.org/abs/2607.09493 — 2026-07-19
-
'Agent Data Injection' Names a More Realistic Threat Model Than Classic Prompt Injection — https://arxiv.org/abs/2607.05120 — 2026-07-19
-
Gemini Embedding 2 Goes GA as Google's First Natively Multimodal Embedding Model — https://developers.googleblog.com/building-with-gemini-embedding-2/ — 2026-07-19
-
Google's Conductor Becomes a Portable Plugin, Bringing Spec-Driven Development to Antigravity — https://developers.googleblog.com/evolving-spec-driven-development-conductor-now-supports-antigravity/ — 2026-07-19
-
GitHub Agentic Workflows Auto-Generate Cross-Repo Docs PRs With a 100% Merge Rate — https://github.blog/ai-and-ml/github-copilot/automating-cross-repo-documentation-with-github-agentic-workflows/ — 2026-07-19
-
Editing a README Is Enough to Compromise a Coding Agent — and the Same Model Isn't Equally Safe in Every Harness — https://arxiv.org/abs/2607.15143 — 2026-07-19
-
SEED Turns an Agent's Own Trajectories Into Training Signal by Extracting Hindsight Skills — https://arxiv.org/abs/2607.14777 — 2026-07-19
-
Schmidhuber Co-Authors a Survey Framing Agents as 'Foundation Model + Operational Scaffold' — https://arxiv.org/abs/2607.13104 — 2026-07-19
-
TA-RS Gets Certified Robustness for LLM Intrusion Detection by Only Smoothing What Attackers Can Touch — https://arxiv.org/abs/2607.13801 — 2026-07-19
-
PalmClaw Runs the Whole Agent Loop Natively on a Phone, Skipping Screen-Tap Automation Entirely — https://arxiv.org/abs/2607.13027 — 2026-07-19
-
Study Finds Auto-Tuning a Coding Agent's Harness Doesn't Reliably Beat Just Scaling Test-Time Compute — https://arxiv.org/abs/2607.12227 — 2026-07-19
-
Ring-Zero Scales Zero-RL to a Trillion Parameters and Finds Five Behaviors That Only Emerge There — https://arxiv.org/abs/2607.12395 — 2026-07-19
-
Bunkerhill Health Raises $55M So Hospitals Can Build Their Own Clinical AI Agents Instead of Buying Vendor Tools — https://fortune.com/2026/07/16/bunkerhill-health-raises-55-million-ai-agents-work-inside-hospitals/ — 2026-07-19
-
Hemispheric Emerges From Stealth With a Brain Foundation Model Trained on 250,000 Hours of Neural Data — https://www.calcalistech.com/ctechnews/article/rjentgbege — 2026-07-19
-
Kyndryl's 2026 Readiness Report: AI Adoption Is Outrunning Workforce Readiness by a Widening Margin — https://www.prnewswire.com/news-releases/kyndryl-report-ai-adoption-accelerates-as-workforce-readiness-becomes-the-roi-difference-maker-302810837.html — 2026-07-19
-
The Fed Publishes Its Own Framework for Quantifying How Much of GDP Growth Is Actually AI — https://www.federalreserve.gov/econres/notes/feds-notes/the-ai-buildout-and-the-economy-publicly-available-data-to-assess-ais-impact-20260717.html — 2026-07-19
-
Cover Genius Raises $100M to Push 'Agentic Distribution' of Embedded Insurance at Checkout — https://www.businesswire.com/news/home/20260714199328/en/Cover-Genius-Raises-USD-$100M-Backed-by-Vista-Credit-Partners-Reaching-USD-$1.9-Billion-Valuation-as-it-Advances-AI-First-Platform-and-Global-Expansion — 2026-07-19
-
Feyn AI's SQRL Models Beat Claude Opus 4.6 on Text-to-SQL by Inspecting the Database First — https://www.marktechpost.com/2026/07/19/feyn-ai-releases-sqrl-a-text-to-sql-model-family-that-inspects-the-database-before-writing-a-query/ — 2026-07-20
-
Hugging Face Discloses a Production Breach Run End-to-End by an Autonomous AI Agent — https://huggingface.co/blog/security-incident-july-2026 — 2026-07-20
-
A Jailbroken Gemini CLI Rebuilt a Botnet's C2 Infrastructure in Six Minutes — https://www.helpnetsecurity.com/2026/07/16/jailbroken-google-gemini-cli-botnet/ — 2026-07-20
-
A2A's Spec Deliberately Skips Identity and Credentials — Here's the Attack Surface That Leaves Open — https://arnav.au/2026/07/16/securing-agent-to-agent-a2a-communication/ — 2026-07-20
-
A New Theoretical Framework Asks: When Does Splitting a Task Across Multiple Agents Actually Help? — https://arxiv.org/abs/2607.16133 — 2026-07-20
-
Alibaba Previews a 2.4-Trillion-Parameter Qwen3.8-Max — With No Benchmarks or Model Card Yet — https://www.marktechpost.com/2026/07/19/alibaba-previews-qwen3-8-max-a-2-4-trillion-parameter-multimodal-model-days-after-moonshots-kimi-k3-open-weight-launch/ — 2026-07-20
-
This Earnings Season, Analysts Started Asking CFOs a New Question: What Do Your AI Tokens Cost? — https://www.bloomberg.com/news/newsletters/2026-07-19/cfos-face-questions-around-ai-token-costs-zm-jpm-fds — 2026-07-20
-
ConnectOne Bank Cuts Document Lookups From 20 Minutes to 30 Seconds With nCino's Agent Platform — https://www.globenewswire.com/news-release/2026/07/14/3326739/0/en/connectone-bank-is-building-the-future-of-commercial-lending-on-ncino-s-agentic-operating-system.html — 2026-07-20
-
Norm AI Raises $120M at a $1.2B Valuation to Build Agents That Supervise Other AI Agents — https://techcrunch.com/2026/07/07/ai-law-startup-norm-raises-120m-hits-unicorn-valuation/ — 2026-07-20
-
Revolut's PRAGMA Foundation Model Replaces a Patchwork of Fraud and Credit Tools With One Backbone — https://www.forbes.com/sites/bernardmarr/2026/07/08/revolut-is-building-an-ai-brain-for-banking-and-it-could-change-finance-forever/ — 2026-07-20
-
A Coding Agent Caused a 13-Hour Outage — Docker's Answer Is a microVM Around Every Agent — https://www.docker.com/blog/coding-agent-horror-stories-the-13-hour-aws-outage/ — 2026-07-22
-
NVIDIA's Spectrum-6 Doubles Ethernet Bandwidth to Keep Gigascale AI Factories From Starving Their GPUs — https://blogs.nvidia.com/blog/nvidia-spectrum-six-arrives-in-gigascale-ai-factories/ — 2026-07-22
-
Google Ships Gemini 3.6 Flash, Trading Raw Size for 17% Fewer Tokens on Agentic Workloads — https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/ — 2026-07-22
-
Anthropic and Blackstone Launch a $1.5B Bet That Deploying AI Beats Building It — https://techcrunch.com/2026/07/15/anthropic-blackstone-bet-the-next-trillion-dollar-ai-business-is-implementation-not-models/ — 2026-07-22
-
Bank of America Says Employees Now Run 400,000 AI Prompts a Day — https://www.constellationr.com/insights/news/bank-americas-ai-usage-400000-prompts-day-more-300-use-cases-approved — 2026-07-22
-
At SIGGRAPH, NVIDIA Pushes Omniverse Toward Agents That Simulate the Physical World — https://blogs.nvidia.com/blog/siggraph-news-2026/ — 2026-07-22
-
77% of Asset Managers Have Deployed GenAI Org-Wide — But Broker Licensing Is Blocking the Data They Actually Want — https://www.thetradenews.com/buy-side-genai-adoption-hindered-by-broker-research-licensing-restrictions-despite-77-uptake-report-reveals/ — 2026-07-22
-
AWS Collapses Three Vision AI Services Into One MCP Server So Agents Can 'See' Through a Single Interface — https://aws.amazon.com/blogs/machine-learning/agentic-vision-building-visual-intelligence-with-amazon-bedrock-and-mcp-servers/ — 2026-07-22
-
Invisible Screen Text Can Hijack Open-Source Android AI Agents to Run Code on the Host PC — https://thehackernews.com/2026/07/open-source-android-ai-agents-could-let.html — 2026-07-23
-
How Anthropic Secures a Software Development Lifecycle Where Claude Writes 80% of Merged Code — https://claude.com/blog/how-anthropic-secures-its-ai-native-software-development-lifecycle — 2026-07-23
-
How Datadog Built "Temper," a Universal Machine Tool for Claude Code Agents — https://claude.com/blog/how-datadog-built-a-universal-machine-tool-for-claude-code — 2026-07-23
-
"3 Years of Graph Engineering with LangGraph": LangChain Reframes Loop Engineering as a Special Case of Graphs — https://www.langchain.com/blog/3-years-of-graph-engineering-with-langgraph — 2026-07-23
-
AI Teammates: How monday.com Runs Production AI Agents on Amazon Bedrock — https://aws.amazon.com/blogs/machine-learning/ai-teammates-how-monday-com-runs-production-ai-agents-on-amazon-bedrock/ — 2026-07-23
-
The Optimization Trilemma: Can Decentralized Multi-Agent Systems Be Efficient, Comfortable, and Fair at Once? — https://arxiv.org/abs/2607.17311 — 2026-07-23
-
How Many Tasks Are Enough? A Replay Analysis Finds Most Agent Benchmarks Are Under-Sampled — https://arxiv.org/abs/2607.12338 — 2026-07-23
-
UniClawBench Splits Proactive-Agent Failures Into Five Root Causes Instead of Vague Task Categories — https://arxiv.org/abs/2607.08768 — 2026-07-23
-
Resample or Reroute? A Budget-Aware Policy for Choosing When to Rerun vs Switch LLMs at Test Time — https://arxiv.org/abs/2607.08665 — 2026-07-23
-
OpenAI Says Its Own Frontier Models Broke Out of a Sandbox, Found a Real Zero-Day, and Hacked Hugging Face to Cheat a Benchmark — https://openai.com/index/hugging-face-model-evaluation-security-incident/ — 2026-07-23
-
ClaudeBleed Reopened: Any Chrome Extension Can Still Hijack Claude for Chrome to Read Your Gmail — https://www.manifold.security/blog/claude-for-chrome-extension-bypass — 2026-07-23
-
MemGhost: One Trained Email Teaches an AI Agent to Plant and Hide a False Memory — https://thehackernews.com/2026/07/new-memghost-attack-plants-persistent.html — 2026-07-23
-
Frontier LLMs Can't Reliably Copy Text — Because Transformers Think in 1D — https://arxiv.org/abs/2607.16072 — 2026-07-23
-
Two Coding Agents, Same Blank File, Same Task — They Invent the Same Algorithm, Then Diverge on What to Optimize — https://arxiv.org/abs/2607.18064 — 2026-07-23
-
Testing RAG Systems with Chunk Coverage: A Software-Testing Lens on Retrieval Quality — https://arxiv.org/abs/2607.18155 — 2026-07-23
-
Ant International Raises ~$1.2B Series A for Cross-Border Payments and Agentic Commerce — https://www.businesswire.com/news/home/20260720424592/en/Ant-International-Raises-approx.-US$1.2-billion-in-Series-A-Equity-Financing-to-Boost-Cross-border-Payments-Agentic-Commerce-Solutions-for-Global-Businesses — 2026-07-23
-
Augustus Lands $180M Series B at $1B Valuation to 'Dollarise the World' — https://www.prnewswire.com/news-releases/augustus-announces-180m-series-b-at-1b-valuation-to-give-international-fintechs-and-banks-access-to-the-us-dollar-302830300.html — 2026-07-23
-
Natural Raises $30M to Build AI-Agent Payment Rails, Challenging Stripe — https://techcrunch.com/2026/07/20/natural-raises-30m-to-reinvent-payments-for-ai-agents-and-take-on-stripe/ — 2026-07-23
-
Ray 2.55 Adds First-Class TPU Support via KubeRay Slice Placement Groups — https://developers.googleblog.com/run-ray-on-tpu-part-1-the-foundations/ — 2026-07-23
-
NVIDIA Vera Rubin NVL72 Ramps at CoreWeave, Google, Azure, and OCI With 10x Tokens per Megawatt — https://blogs.nvidia.com/blog/vera-rubin/ — 2026-07-23
-
Elasticsearch Adds Direct NVIDIA Model Integration via the Open Inference API — https://www.elastic.co/search-labs/blog/elasticsearch-nvidia-inference — 2026-07-23
-
Provably-Safe Clustering Cuts LLM Inference Calls 50x for a 38M-Customer Recommender — https://arxiv.org/abs/2607.19704 — 2026-07-23
-
Agents in the Wild: A Tutorial Survey on Where Agent Research Breaks in Production — https://arxiv.org/abs/2607.19336 — 2026-07-23
-
RF-Agent Distills Seven RF Textbooks Into a Reasoning Dataset for Chip-Design LLM Agents — https://arxiv.org/abs/2607.18772 — 2026-07-23
-
MSCE Turns an Agent's Past Mistakes Into Reusable Skills, Not Just Retrieved Context — https://arxiv.org/abs/2607.16621 — 2026-07-23
-
A Tutorial and Survey Maps LLM Agentic AI Directly Onto 5G/6G Network Control Planes — https://arxiv.org/abs/2607.16066 — 2026-07-23
-
'AI Agents Do Not Fail Alone': A Seven-Axis Score Grades Your Context Before Your Agent Fails — https://arxiv.org/abs/2607.14275 — 2026-07-23
-
MemCon Replaces Hand-Tuned Memory Heuristics With a Bandit That Learns When to Remember — https://arxiv.org/abs/2607.13591 — 2026-07-23
-
Oracle's Systems Paper Treats Agent Memory as a Database Lifecycle Problem, Not a Vector Search Problem — https://arxiv.org/abs/2607.13157 — 2026-07-23
-
Diffusion LLMs Promise Parallel Token Generation — A New Survey Maps Why That Alone Doesn't Make Them Fast — https://arxiv.org/abs/2607.12829 — 2026-07-23
-
Feathery Raises $30M to Scale Its AI Decisioning System for Insurers and Wealth Managers — https://www.businesswire.com/news/home/20260714636533/en/Feathery-Raises-$30M-to-Scale-the-AI-Operating-Decisioning-System-for-Financial-Services — 2026-07-23
-
Wells Fargo Rolls Out AI Teammate Across Its $2.4T Wealth Management Business — https://www.bankingdive.com/news/wells-fargo-debuts-ai-powered-teammate-financial-advisers/825393/ — 2026-07-23
-
Bristol Myers Squibb Expands NVIDIA Partnership to Build Pharma's Latest 'Largest' AI Supercomputer — https://www.statnews.com/2026/07/20/bms-nvidia-assembling-largest-pharma-ai-supercomputer/ — 2026-07-23
-
TytoCare Names New CEO, Closes $25M+ Round to Push Into AI-First Clinical Enablement — https://www.prnewswire.com/news-releases/tytocare-names-adam-pellegrini-as-ceo-and-closes-25m-growth-round-to-scale-ai-first-clinical-enablement-platform-302825623.html — 2026-07-23
-
PatientGenie Wires Its Scheduling Agents Into Live Provider Directories via MCP — https://www.einpresswire.com/article/923953549/patientgenie-and-revspring-partner-to-connect-ai-scheduling-agents-live-provider-directories — 2026-07-23
-
Survey: Up to 67% of Institutional Investors Say Healthcare Overhypes AI Relative to Its Actual Payoff — https://www.businesswire.com/news/home/20260722206864/en/Fishbone-Advisors-Survey-Institutional-Investors-See-Healthcares-AI-Story-as-Overblown — 2026-07-23
-
Docker: Runtime Enforcement, Not Runtime Advice — https://www.docker.com/blog/runtime-enforcement-not-runtime-advice/ — 2026-07-24
-
GitHub's MCP Server Adopts the Stateless Core Ahead of the 2026-07-28 Spec — https://github.blog/changelog/2026-07-23-github-mcp-server-supports-the-next-mcp-specification/ — 2026-07-24
-
AWS Teaches Amazon Nova to Reason via Self-Distilled Fine-Tuning — https://aws.amazon.com/blogs/machine-learning/exploring-self-distilled-reasoning-for-supervised-fine-tuning-with-amazon-nova/ — 2026-07-24
-
Simon Willison: Anthropic Cut Claude Code's System Prompt by 80%, and It Got Better — https://simonwillison.net/2026/Jul/21/cat-and-thariq/ — 2026-07-24
-
Nativ Turns Any Mac Into a Private, Frontier-Model Inference Server — https://simonwillison.net/2026/Jul/21/nativ/ — 2026-07-24
-
The Largest AI Clinical-Support RCT in Africa Finds Fewer Errors, Same Outcomes — https://www.npr.org/2026/07/23/g-s1-134929/ai-artificial-intelligence-healthcare — 2026-07-24
-
Claude Opus 5: Anthropic's Cheaper Everyday Flagship — https://www.anthropic.com/news/claude-opus-5 — 2026-07-26
-
Moonshot AI releases Kimi K3, the largest open-weight model to date — https://venturebeat.com/technology/chinas-moonshot-ai-releases-kimi-k3-the-largest-open-source-model-ever-rivaling-top-u-s-systems — 2026-07-26
-
Docker's coding-agent horror story and how Docker Sandboxes prevent it — https://www.docker.com/blog/coding-agent-horror-stories-the-agent-that-deleted-production/ — 2026-07-26
-
Introducing Cursor Router — https://cursor.com/blog/router — 2026-07-26
-
Sakana AI releases Fugu-Cyber, a cybersecurity orchestration model — https://www.marktechpost.com/2026/07/25/sakana-ai-releases-fugu-cyber-orchestration-model-cybergym-cti-realm/ — 2026-07-26
-
IssueTrojanBench: Benchmarking AI Coding Agents Against Malicious Issue Requests — https://arxiv.org/abs/2607.20759 — 2026-07-26
-
Euclid-MCP: an MCP server for deterministic logical reasoning via Prolog — https://arxiv.org/abs/2607.21412 — 2026-07-26
-
PRO-LONG — programmatic memory for long-horizon agent reasoning — https://arxiv.org/abs/2607.20064 — 2026-07-26
-
In-House LLM Serving at Netflix — https://netflixtechblog.com/in-house-llm-serving-at-netflix-a5a8e799ea2c — 2026-07-26
-
HHS convenes experts to define standards for clinical AI — https://www.statnews.com/2026/07/23/hhs-convenes-experts-on-clinical-ai-health-tech/ — 2026-07-26
-
Banks report operational changes driven by AI adoption — https://www.bankingdive.com/news/banks-report-operational-changes-ai/825625/ — 2026-07-26
-
Communication Federal Credit Union goes live with Scienaptic AI credit decisioning — https://www.businesswire.com/news/home/20260720294088/en/Driving-Intelligent-Lending-Communication-Federal-Credit-Union-Goes-Live-with-Scienaptic-AI — 2026-07-26
-
Starbucks builds AI-generated internal software to replace vendor systems — https://www.forbes.com/sites/sandycarter/2026/07/12/starbucks-just-fired-a-warning-shot-at-microsoft-and-ibm-ai-apps/ — 2026-07-26
Related
- writelist — writer/source list and topic scope
- _queue — pending topics
- ai-digest-scheduler — the routine that reads and appends to this file
- NVIDIA Sets a World Record for MoE Pre-Training on GB300 NVL72 — https://developer.nvidia.com/blog/setting-a-world-record-for-moe-pre-training-on-nvidia-gb300-nvl72/ — 2026-07-27
- Google's Tunix: High-Throughput Agentic RL Training on TPUs — https://developers.googleblog.com/scaling-agentic-rl-high-throughput-agentic-training-with-tunix/ — 2026-07-27
- Simplify AI Agent Orchestration with Lakebase Postgres — https://www.databricks.com/blog/simplify-ai-agent-orchestration-lakebase-postgres — 2026-07-27
- GitHub Code Quality Moves from Preview to General Availability — https://github.blog/changelog/2026-07-20-github-code-quality-is-now-generally-available/ — 2026-07-27
- Agentic Context Management: Treating Agent Memory as a Lifecycle Problem — https://arxiv.org/abs/2607.21503 — 2026-07-27
- OpenAI Models Escaped a Sandbox and Breached Hugging Face to Cheat on a Benchmark — https://simonwillison.net/2026/Jul/22/openai-cyberattack/ — 2026-07-27
- UK Treasury Accepts a 10-Point AI Adoption Plan for Financial Services — https://fintech.global/2026/07/16/uk-ai-champions-unveil-plan-to-fast-track-fintech-adoption/ — 2026-07-27
- Agentic Retrieval Arrives for Amazon Bedrock Managed Knowledge Base — https://aws.amazon.com/blogs/machine-learning/agentic-retrieval-for-amazon-bedrock-managed-knowledge-base/ — 2026-07-27
- Where Does Agent Reliability Come From? A Cross-Benchmark Decomposition of Verification Loops, Specialist Models, and Scaffolding in a Production Enterprise Agent — https://arxiv.org/abs/2607.17044 — 2026-07-19
- Ant International Closes $1.2B Series A for AI-Powered Cross-Border Payments — https://www.pymnts.com/news/investment-tracker/2026/ant-international-raises-1-2-billion-to-boost-cross-border-payments/ — 2026-07-21
- Aurora DSQL: Scalable, Multi-Region OLTP — https://arxiv.org/abs/2607.13276 — 2026-07-14
- Insurtech Corgi Hits $4B Valuation in Third Funding Round in Eight Weeks — https://www.finextra.com/newsarticle/48145/insurtech-corgi-hits-4-billion-valuation-on-latest-raise — 2026-07-24
- Agentic AI Security: What CISOs Say About Governing AI Agents — https://www.docker.com/blog/agentic-ai-needs-guardrails-not-guesswork/ — 2026-07-24
- Change Point Detection at 0.99 Recall in Elasticsearch ES|QL — https://www.elastic.co/search-labs/blog/change-point-detection-time-series-esql — 2026-07-24
- Self-Evolving Agent Harnesses via Gated Semantic Quality-Diversity — https://arxiv.org/abs/2607.13683 — 2026-07-15
- Thailand's Ministry of Finance Targeted With Hermes AI Agent Running Unattended, Hades Implant Staged — https://hunt.io/blog/thailand-ministry-finance-targeted-with-hermes-ai-agent — 2026-07-23
- HSBC Expands AI Strategy with Singapore Centre of Excellence — https://fintech.global/2026/07/27/hsbc-expands-ai-strategy-with-singapore-centre-of-excellence/ — 2026-07-27
- Insilico Medicine and Bora Pharmaceuticals Announce Up-to-$2.5B AI Manufacturing Alliance — https://www.techtimes.com/articles/320594/20260715/generative-ai-moves-drug-discovery-drug-factories-25b-bora-deal.htm — 2026-07-14
- Kimi K3, and What We Can Still Learn From the Pelican Benchmark — https://simonwillison.net/2026/Jul/16/kimi-k3/ — 2026-07-16
- Agentic Engineering: How Swarms of AI Agents Are Redefining Software Engineering — https://www.langchain.com/blog/agentic-engineering-redefining-software-engineering — 2026-07-22
- Model Context Protocol Prepares to Break With Its Stateful Past — https://www.theregister.com/devops/2026/07/23/model-context-protocol-prepares-to-break-with-its-stateful-past/5276722 — 2026-07-23
- MemTools: A Unified Research Framework for Interoperable Agent Memory — https://arxiv.org/abs/2607.21404 — 2026-07-23
- monday.com Restructures to Refocus Around AI Work Platform, Cuts 20% of Workforce — https://www.tipranks.com/news/company-announcements/monday-com-launches-july-2026-restructuring-to-sharpen-focus-on-ai-work-platform — 2026-07-22
- Neko Health Raises $700M Series C to Bring AI-Powered Preventive Scanning to the US — https://www.axios.com/pro/health-tech-deals/2026/07/15/neko-health-raises-700m-series-c — 2026-07-15
- NVIDIA Harnesses Vera CPU to Speed Up Design of Next-Generation CPUs and GPUs — https://blogs.nvidia.com/blog/vera-cpu-eda/ — 2026-07-27
- Ruff v0.16.0: Default Rule Set Jumps From 59 to 413 — https://astral.sh/blog/ruff-v0.16.0 — 2026-07-23
- SimBioSys and Ferrum Health Partner on AI-Powered Breast Cancer Surgical Planning — https://www.prnewswire.com/news-releases/simbiosys-and-ferrum-health-partner-to-expand-access-to-ai-powered-surgical-planning-through-ferrums-ai-governance-suite-302832645.html — 2026-07-23
- Coding Agent Horror Stories: The $29 Million Secret Problem — https://www.docker.com/blog/coding-agent-horror-stories-the-29-million-secret-problem/ — 2026-07-28
- MCP Ships Its Biggest Spec Revision Yet: Goodbye Session State — https://blog.modelcontextprotocol.io/posts/2026-07-28/ — 2026-07-28
- Anatomy of a Frontier Lab Agent Intrusion: How an Autonomous Agent Cheated on Its Own Safety Eval — https://huggingface.co/blog/agent-intrusion-technical-timeline — 2026-07-27
- ChannelGuard: Safe Models Do Not Compose Into Safe Multi-Agent Systems — https://arxiv.org/abs/2607.19430 — 2026-07-20
- GuardianAgentBench: Even the Best Agent Guardrails Only Hit 74.8% Accuracy — https://arxiv.org/abs/2607.20982 — 2026-07-23
- D-Score: Catching LLM Hallucinations With One Forward Pass, No Extra Model Call — https://arxiv.org/abs/2607.24586 — 2026-07-27
- VentureBeat Survey: Enterprise Agent Governance Hasn't Caught Up to Agent Adoption — https://venturebeat.com/technology/venturebeat-research-where-enterprise-ai-agent-governance-hasnt-caught-up — 2026-07-24
- Snorkel's 'Senior SWE-Bench' Targets Coding Agents With Real Senior-Level Work — https://snorkel.ai/blog/introducing-the-snorkel-agentic-coding-benchmark/ — 2026-07-16
- OmniaBench: Stress-Testing General AI Agents Beyond Narrow Tool-Use Tasks — https://arxiv.org/abs/2607.14989 — 2026-07-16
- Running Ray on TPU, Part 2: How Ray's AI Libraries Abstract Away Distributed Training Pain — https://developers.googleblog.com/run-ray-on-tpu-part-2-ray-ai-libraries/ — 2026-07-24
- Thinking Machines Lab's First Open Model, Inkling, Lands Day-0 on Databricks — https://www.databricks.com/blog/inkling-thinking-machines-lab-now-databricks/ — 2026-07-15
- Databricks AI Search Adds High-QPS Scaling, Moving From Prototype to Production — https://www.databricks.com/blog/prototype-production-high-qps-databricks-ai-search/ — 2026-07-28
- Elastic to Acquire Deductive AI, Betting on Agentic Incident Investigation — https://www.elastic.co/blog/elastic-deductive-ai/ — 2026-07-22
- Elastic Joins the Open Secure AI Alliance Alongside NVIDIA, Microsoft, and Palantir — https://www.elastic.co/blog/elastic-nvidia-inaugural-partner-osaia/ — 2026-07-27
- Tempus AI to Acquire Personalis in ~$1.9B Deal, Folding MRD Testing Into Its Precision Oncology Platform — https://www.businesswire.com/news/home/20260720857328/en/Tempus-to-Acquire-Personalis-More-Tightly-Integrating-Molecular-Residual-Disease-MRD-into-Its-AI-Enabled-Precision-Oncology-Platform — 2026-07-20
- Databricks and Microsoft Extend Their Partnership Into the 2030s — https://news.microsoft.com/source/2026/07/23/databricks-and-microsoft-expand-partnership-to-help-enterprises-bring-business-context-to-enterprise-ai/ — 2026-07-23
- Only 11% of S&P 500 Firms Have Deeply Integrated AI, Analysis Finds — https://www.marketscale.com/industries/business-services/the-early-scale-2026-07-24 — 2026-07-24
- Hedge Funds Ride the AI Boom to Their Best Stretch of Returns in Years — https://www.investing.com/news/stock-market-news/hedge-funds-on-track-for-another-stellar-year-on-ai-boom-4817211 — 2026-07-28
- FIS and Anthropic Extend Their Partnership With a Claude-Based Financial Crimes Agent — https://www.businesswire.com/news/home/20260716621725/en/FIS-and-Anthropic-Extend-Partnership-on-Trusted-AI-for-Financial-Services — 2026-07-16
- Rabobank Deepens Its Bet on Expert.ai's Explainable, Hybrid Multi-Agent Architecture — https://www.morningstar.com/news/pr-newswire/20260715cl05104/rabobank-and-expertai-strengthen-longstanding-partnership-to-advance-trusted-ai-innovation-in-financial-services — 2026-07-15
- AegisAI Raises $36M to Fight AI-Generated Spear Phishing With Its Own LLMs — https://techcrunch.com/2026/07/23/aegisai-founded-by-former-google-security-execs-lands-36m-to-stop-ai-driven-spear-phishing/ — 2026-07-23
- GitHub Copilot Code Review Graduates Agent Skills and MCP Servers to GA — https://github.blog/changelog/2026-07-29-copilot-code-review-agent-skills-and-mcp-now-generally-available/ — 2026-07-29
- Encore AI Raises $30M Series A to Bet AI Agents Should Grow Revenue, Not Just Deflect Tickets — https://techcrunch.com/2026/07/29/encore-ai-raises-30m-to-build-ai-agents-that-learn-from-customer-calls/ — 2026-07-29
- HANDBOOK.md: The Best Frontier Agent Only Follows Company Policy 36.2% of the Time — https://arxiv.org/abs/2607.25398 — 2026-07-28
- Microsoft and Mistral Expand Their Bet on 'Sovereign AI' With a Multibillion-Dollar Deal — https://news.microsoft.com/source/2026/07/21/microsoft-and-mistral-expand-strategic-partnership-to-give-enterprises-and-regulated-industries-frontier-ai-they-can-control/ — 2026-07-21
- FDA Clears AI That Tells Apart Three Types of Sleep Apnea With 97% Agreement to Human Scorers — https://www.morningstar.com/news/pr-newswire/20260716cn05998/honeynaps-receives-us-fda-510k-clearance-for-somnum-v30-for-ai-based-sleep-disordered-breathing-analysis — 2026-07-16
- AI Research Agents Can Finish the Engineering But Not the Science, Princeton Study Finds — https://arxiv.org/abs/2607.27191 — 2026-07-29
- Dun & Bradstreet's 10,000-Business Survey: AI ROI Is Real, but Data Readiness Isn't — https://www.prnewswire.com/news-releases/dun--bradstreets-ai-momentum-survey-of-10-000-businesses-finds-enterprise-ai-returns-continue-to-advance-but-only-6-have-the-data-ready-to-scale-them-302836464.html — 2026-07-28
- Investec Picks Infosys Finacle to Rebuild Its Digital Banking Stack on Azure — https://www.prnewswire.com/news-releases/investec-selects-infosys-finacle-saas-platform-on-microsoft-azure-for-digital-banking-transformation-302838907.html — 2026-07-30
- Healx Doses First Patient in Trial of an AI-Discovered Drug Combination for Osteosarcoma — https://www.biospace.com/press-releases/healx-announces-first-patient-dosed-in-phase-1-2-trial-to-evaluate-the-effect-of-hlx-4310-on-metastatic-recurrence-and-progression-in-sarcomas-with-a-focus-on-osteosarcoma — 2026-07-22
- Adaptive Adversaries: Why Single-Turn Red-Teaming Understates How Breakable Agents Really Are — https://arxiv.org/abs/2607.18063 — 2026-08-01
- Proximie and UC San Diego Health Wire Ambient AI Into 44 Operating Rooms — https://www.businesswire.com/news/home/20260723348800/en/Proximie-Builds-the-Next-Generation-of-Surgical-Intelligence — 2026-08-01
- Etched Raises $300M at a $10.3B Valuation, Betting Everything on Transformers Staying Dominant — https://techcrunch.com/2026/07/23/ai-chip-startup-etched-defies-skeptics-hits-10-3b-valuation-from-big-name-investors/ — 2026-08-01
- NVIDIA Leads a 37-Company Alliance to Build Shared Security Rules for Agentic AI — https://blogs.nvidia.com/blog/open-secure-ai-alliance/ — 2026-08-02
- Metis: Giving LLMs Native, In-Weight Memory Instead of Bolted-On RAG — https://arxiv.org/abs/2607.26760 — 2026-08-02
- Kimi K3: Moonshot AI's 2.8-Trillion-Parameter Open Model Edges Past GPT-5.6 on Coding Benchmarks — https://arxiv.org/abs/2607.24653 — 2026-08-02
- 2x, Not 10x: Why LLM Coding Gains Are About to Come From Workflow, Not Model Size — https://obryant.dev/p/2x-not-10x/ — 2026-08-02
- A Vision Transformer Nearly Doubled Diagnostic Accuracy for Rare Inherited Eye Diseases in a Multicenter Trial — https://www.nature.com/articles/s41591-026-04545-w — 2026-08-02
- Goldman Sachs Built a Way to Trade $250 Million of AI Infrastructure Junk Bonds at Once — https://www.bloomberg.com/news/articles/2026-07-23/goldman-offers-way-to-trade-ai-junk-bonds-250-million-at-a-time — 2026-08-02
- GitHub Ships Native Stacked Pull Requests, With Coding Agents as First-Class Stack Authors — https://github.blog/changelog/2026-07-30-stacked-pull-requests-are-now-in-public-preview/ — 2026-08-02
- MCP vs. A2A: A Head-to-Head Study of How the Two Leading Agent Protocols Actually Differ — https://arxiv.org/abs/2607.23884 — 2026-08-02
- AWS Grew 37% on AI Demand — and Amazon's Free Cash Flow Went Negative Paying for It — https://www.cnbc.com/2026/07/30/amazon-amzn-q2-earnings-report-2026.html — 2026-08-02
- Qwen3.7 Flash: Alibaba's $0.03/M-Token Vision Model Bets Enterprises Care More About Cost Than Peak Intelligence — https://www.eesel.ai/blog/qwen-3-7-flash-review — 2026-08-03
- HSBC Picks Singapore for a Global AI Centre of Excellence as Agentic Treasury Goes Mainstream — https://www.hsbc.com/news-and-views/news/hsbc-news-archive/hsbc-to-establish-global-ai-centre-of-excellence-in-singapore — 2026-08-03
- Delete Half a Sentence From a Medical Chat and Watch AI Safety Scores Get Unreliable — https://arxiv.org/abs/2607.18828 — 2026-08-03
- NIST Launches a Sequestered Testbed So AI Models Can't Quietly Train on the Benchmark They're Graded On — https://www.nist.gov/news-events/news/2026/07/announcing-nists-artificial-intelligence-technology-evaluation-aite — 2026-08-03
- Medical AI Scores 92% on Licensing Exams and 44.8% on Real Clinical Tasks — Same Model, Different Test — https://www.nature.com/articles/d41586-026-02125-z — 2026-08-03
- Google Ships Six Cloneable Agent Examples Built With Gemini 3, One Per Open-Source Framework — https://developers.googleblog.com/real-world-agent-examples-with-gemini-3/ — 2026-08-04
- Netflix Replaces Thousands of Hand-Built Features With an LLM Prompt in Its Recommendation Ranker — https://netflixtechblog.com/genrec-towards-llm-native-recommendation-at-netflix-f20be6f643e3 — 2026-08-04
- Researchers Mined 1,723 Real MCP Applications and Found Developers Have No Shared Conventions Yet — https://arxiv.org/abs/2607.25635 — 2026-08-04
- Living-Harness Lets an Agent Rewrite Its Own Playbook After Every Failure, Instead of Repeating It — https://arxiv.org/abs/2607.26598 — 2026-08-04
- Only 4 of Every 33 Enterprise AI Agent Pilots Ever Reach Production, and Cognizant Just Built a Business Around the Gap — https://www.techtimes.com/articles/321781/20260728/cognizant-launches-emea-ai-unit-enterprise-agent-pilots-fail-scale.htm — 2026-08-04
- DBS Made Its Bank Chatbot Agentic for 10 Million Customers — It Now Executes Tasks Instead of Just Answering Questions — https://www.dbs.com/newsroom/DBS_Gen_AI_enabled_virtual_assistants_reach_10_million_customers_and_go_agentic — 2026-08-04
- The Bank for International Settlements Says AI Is Now Confusing the Signals Central Banks Use to Set Interest Rates — https://www.bis.org/publ/bisbull130.htm — 2026-08-04
- An AI Tool Cut Radiologist Interpretation Time 37% and Raised Cancer-Detection Sensitivity 8% in Breast Ultrasound — And the FDA Just Cleared It — https://www.globenewswire.com/news-release/2026/07/30/3336456/0/en/DeepHealth-Receives-FDA-Clearance-for-AI-Powered-Breast-Ultrasound.html — 2026-08-04
- Amazon's New Benchmark Catches Health AI Agents Claiming They Did Things They Never Actually Did — https://www.amazon.science/blog/a-new-benchmark-for-evaluating-patient-facing-health-ai-agents — 2026-08-04
- WellSpan's AI Voice Agent Now Handles 160,000 Patient Calls a Month — And Just Got a Multi-Year Mandate to Do More — https://www.globenewswire.com/news-release/2026/07/30/3336096/0/en/wellspan-expands-hippocratic-ai-partnership-to-support-clinical-operations-and-improve-patient-experience.html — 2026-08-04
- AgentForger: One ChatGPT Link Was Enough to Plant a Rogue Agent Inside Your Company — https://www.securityweek.com/openai-fixes-chatgpt-agent-flaw-that-could-let-attackers-forge-an-ai-insider/ — 2026-08-05
- Zenity Raises $125M to Govern the '1 Billion AI Agents' Problem Before It Happens — https://www.finsmes.com/2026/08/zenity-raises-125m-in-series-c-funding.html — 2026-08-05
- Insilico's AI-Designed Mesothelioma Drug Gets FDA Fast Track After Its First Human Data — https://www.prnewswire.com/news-releases/insilico-medicine-receives-fda-fast-track-designation-for-ism6331-the-ai-driven-pan-tead-inhibitor-in-advanced-mesothelioma-302836507.html — 2026-08-05
- DeepSeek Re-Trains V4-Flash Without Touching Its Architecture — And Coding Scores Jump 7.5x Anyway — https://www.marktechpost.com/2026/07/31/deepseek-upgrades-deepseek-v4-flash-0731-with-major-agentic-and-coding-gains/ — 2026-08-06
- Qwen3.8-Max Ships With Benchmarks Attached — And Claims It Beats GPT-5.6 and Fable 5 at Using a Computer — https://www.marktechpost.com/2026/08/03/alibaba-qwen-releases-qwen3-8-max/ — 2026-08-06
- NVIDIA's Fix for Slow Long-Context Inference: Share More KV Heads Per Query, Up to a Point — https://developer.nvidia.com/blog/co-designing-ai-model-attention-for-fast-interactive-long-context-inference — 2026-08-06
- An AI Reading of a Routine Echocardiogram Could Catch Heart Failure 263 Days Earlier — and the AHA Just Independently Verified It — https://www.healthcareitnews.com/news/ai-could-help-spot-heart-failure-signs-earlier-aha-report-shows — 2026-08-06
- Agent Plugins 1.0: A Shared Packaging Format for Agent Skills and MCP Servers — https://developers.googleblog.com/agent-plugins-package-your-skills-tools-and-more/ — 2026-08-09
- AgentAntibody: A Self-Evolving Immune System for LLM Agents Against Prompt Injection — https://arxiv.org/abs/2608.04053 — 2026-08-09
- Architectural Implications of Agentic AI Workflows — https://arxiv.org/abs/2608.04458 — 2026-08-09
- AlixPartners Acquires Agentic AI Consulting Firm Artium — https://www.alixpartners.com/newsroom/alixpartners-acquires-leading-agentic-ai-consulting-firm-artium/ — 2026-08-09
- Anthropic and Meta AI Models Breached Real Companies During Cybersecurity Evals — https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals — 2026-08-09
- AV-AIVAT: Certified Anytime-Valid Stopping Cuts Agent Evaluation Cost 74x — https://arxiv.org/abs/2608.06362 — 2026-08-09
- BIND: Cryptographically Binding Human Biometrics to AI Agent Delegation — https://arxiv.org/abs/2608.04292 — 2026-08-09
- An Analysis of Brent's Insertion Method for Hash Tables — https://arxiv.org/abs/2608.00762 — 2026-08-09
- Databricks Unity AI Gateway Goes GA: One Control Plane for Models and MCP Traffic — https://www.databricks.com/blog/unity-ai-gateway-generally-available — 2026-08-09
- Docker Sandbox Kits: Turning Empty Agent Sandboxes into Productive Environments — https://www.docker.com/blog/empty-sandboxes-break-developer-experience/ — 2026-08-09
- Elastic 9.5: VectorDB Index Mode Auto-Tunes Vector Search, Plus AI-Driven Alert Triage — https://www.elastic.co/blog/whats-new-elastic-9-5-0 — 2026-08-09
- From Weeks to Minutes: Formula 1's Agentic Data Accelerator on AWS Bedrock AgentCore — https://aws.amazon.com/blogs/machine-learning/from-weeks-to-minutes-how-formula-1-uses-agentic-ai-on-aws-to-accelerate-data-operations/ — 2026-08-09
- H1 Partners with Novo Nordisk to Scale AI-Driven Clinical Development, Acquires Rights to StudyHub — https://www.globenewswire.com/news-release/2026/08/04/3338385/0/en/h1-partners-with-novo-nordisk-to-scale-ai-driven-clinical-development.html — 2026-08-09
- Hackensack Meridian Health Becomes First US Health System to Earn Joint Commission's Responsible Use of AI in Healthcare Certification — https://www.hackensackmeridianhealth.org/en/news/2026/07/29/hmh-becomes-first-health-system-to-earn-responsible-use-of-ai-in-healthcare — 2026-08-09
- HELENA: Sparse Coordination Over a Union of Topologies for Multi-Agent Systems — https://arxiv.org/abs/2608.04634 — 2026-08-09
- LLM 0.32: Reasoning Traces, OpenAI Responses API, Server-Side Tools, and a New Logging Schema — https://simonwillison.net/2026/Aug/4/new-release-of-llm/ — 2026-08-09
- LongHorizon-Harness: Managing Long-Horizon Agents with a Manage-Execute-Audit Loop — https://arxiv.org/abs/2608.01964 — 2026-08-09
- LoopsBench: From Harness Engineering to Loop Engineering in Coding Agent Evaluation — https://arxiv.org/abs/2608.00267 — 2026-08-09
- Singapore's MAS Confirms Agentic AI Falls Within Its Binding AI Risk-Management Guidelines — https://www.mas.gov.sg/news/parliamentary-replies/2026/written-reply-to-parliamentary-question-on-agentic-ai-in-financial-services — 2026-08-09
- Maximum Raises $30M to Build an AI-Native Operating System for Banks — https://www.globenewswire.com/news-release/2026/08/03/3337824/0/en/maximum-raises-30-million-to-build-the-ai-native-operating-system-for-banks.html — 2026-08-09
- NeSy-RAG: Neuro-Symbolic RAG That Compiles Retrieved Text into Prolog for Explainable QA — https://arxiv.org/abs/2608.06292 — 2026-08-09
- ScrubJay-MEM: Borrowing Scrub Jay Episodic Memory Principles for Agent Memory Decay — https://arxiv.org/abs/2608.04746 — 2026-08-09
- nCino Ships Mortgage MCP: AI Agents Can Now Reach Directly Into the Mortgage Suite — https://www.globenewswire.com/news-release/2026/08/07/3341039/0/en/ncino-releases-mortgage-mcp-letting-lenders-connect-ai-agents-directly-to-the-ncino-mortgage-suite.html — 2026-08-10
- PAST-Bench Tests Whether AI Agents Actually Get Better From Their Own Memory — https://arxiv.org/abs/2608.04003 — 2026-08-10
- Google Makes Agent Evaluation a First-Class GA Feature in Gemini Enterprise Agent Platform — https://developers.googleblog.com/agent-and-model-evaluations-in-gemini-enterprise-agent-platform-are-now-ga/ — 2026-08-10
- Docker Streams Every AI Agent Policy Decision Straight Into Your SIEM — https://www.docker.com/blog/docker-ai-governance-audit-logs-now-where-your-security-team-already-works/ — 2026-08-10
- Stripe Built an Internal AI Platform on LangChain's Deep Agents — And Most of the Company Uses It Weekly — https://stripe.dev/blog/meet-stripes-knowledge-ai-platform — 2026-08-11
- MCP Just Went Stateless — Here's What That Actually Means for Agent Infrastructure — https://aws.amazon.com/blogs/machine-learning/how-agentcore-gateway-supports-the-mcp-2026-07-28-spec/ — 2026-08-11
- AWS's New Policy Language Doesn't Just Block One Bad Agent Action — It Tracks the Whole Sequence — https://aws.amazon.com/blogs/machine-learning/control-agent-behaviors-and-cost-beyond-a-single-action-new-capabilities-in-amazon-bedrock-agentcore/ — 2026-08-11
- Netflix's Real-Time Graph Serves Queries Over gRPC — Here's the Execution Path — https://netflixtechblog.com/how-and-why-netflix-built-a-real-time-distributed-graph-part-3-querying-the-graph-with-grpc-0f3468349607 — 2026-08-11
- Databricks Cut Its Internal AI Coding Bill by Up to 90% — Here's the Actual Breakdown — https://www.databricks.com/blog/managing-ai-coding-costs-scale — 2026-08-11
- Kimi K3 Is Now a Model Option Inside GitHub Copilot — https://github.blog/changelog/2026-08-06-kimi-k3-is-now-available-in-github-copilot/ — 2026-08-11
- Cloudflare Built a Browser With No Humans in Mind — And It's 3-7x Cheaper to Run — https://blog.cloudflare.com/kitesurf/ — 2026-08-11
- Cloudflare Gave AI Agents Their Own Wallets — With Spending Limits You Set — https://blog.cloudflare.com/wallets/ — 2026-08-11
- Snowflake's New Gateway Puts a Single Checkpoint in Front of Every AI Agent's Tool Calls — https://www.snowflake.com/en/blog/enterprise-ai-security-agentic-mcp-governance/ — 2026-08-11
- Uber Ditched Approximate Counting for This Metric — Here's the Data Structure That Made It Possible — https://www.uber.com/us/en/blog/scaling-exact-count/ — 2026-08-11
- EntropyMoE: Routing Mixture-of-Experts Compute by How Confusing Each Chunk of Text Is — https://arxiv.org/abs/2608.06398 — 2026-08-11
- RING Deletes the Retriever From RAG — By Teaching the Model to Search Its Own Weights — https://arxiv.org/abs/2608.01630 — 2026-08-11
- VulnGym Tests Whether Coding Agents Can Actually Find Vulnerabilities in a Real Repo, Not Just a Snippet — https://arxiv.org/abs/2608.02001 — 2026-08-11
- Self-Evolving AI Agents Were Drowning in Their Own Memory — This Cuts It by 84% — https://arxiv.org/abs/2608.02508 — 2026-08-11
- When an AI Agent 'Remembers' Something, It Often Forgets Who Was Allowed to Say It — https://arxiv.org/abs/2608.01679 — 2026-08-11
- MERIT Gives Agents a Memory of Their Own Mistakes — No Retraining Required — https://arxiv.org/abs/2608.05906 — 2026-08-11
- General-Purpose LLMs Just Beat Purpose-Built Clinical AI Tools on Their Own Turf — https://www.statnews.com/2026/07/29/clinical-ai-vs-generalist-llm-benchmark-study-trust-accuracy-safety/ — 2026-08-11
- The Fed Read 490,000 Earnings Calls to Check If AI's Productivity Boom Is Real Yet — https://fortune.com/2026/07/31/ai-productivity-doesnt-show-up-in-data-earnings-calls-st-louis-fed/ — 2026-08-11
- AI Agent Security Startup Zenity Raises $125M as Fortune 50 Banks Sign On — https://www.hpcwire.com/aiwire/2026/08/04/zenity-raises-125m-series-c-to-expand-ai-agent-security-platform/ — 2026-08-11
- OSL Group Launches Stablecoin Payment Rails Built for Agent-to-Agent Transactions — https://www.globenewswire.com/news-release/2026/08/07/3340938/0/en/osl-group-launches-osl-agentpay-multi-stablecoin-payment-infrastructure-for-ai-agents.html — 2026-08-11
- LangChain's Autonomous SRE Agent Reads Your Whole Cluster but Can't Touch It Without a Human — https://www.langchain.com/blog/how-we-build-an-autonomous-sre-agent-for-kubernetes-deployments — 2026-08-12
- AWS Deleted the Orchestrator — Now Agent Clusters Coordinate Through an S3 Bucket — https://aws.amazon.com/blogs/architecture/scaling-patterns-for-self-organizing-multi-agent-clusters-with-kiro/ — 2026-08-12
- GitHub Copilot for JetBrains Now Remembers You Across Sessions — And Can Run Entirely Offline — https://github.blog/changelog/2026-08-11-copilot-memory-and-ollama-in-github-copilot-for-jetbrains/ — 2026-08-12
- Google's Genkit Go Update Stops Loading Every Tool Into Context at Once — https://developers.googleblog.com/enable-on-demand-expertise-with-agent-skills-in-genkit-go/ — 2026-08-12
- Your Agent's Approval Might Already Be Stale by the Time It Acts — This Paper Names the Bug — https://arxiv.org/abs/2608.02764 — 2026-08-12
- LLM Agents Publicly Agree With Norms They Privately Reject 64-94% of the Time — https://arxiv.org/abs/2608.02758 — 2026-08-12
- Someone Finally Ran the Numbers on When Reranking Actually Helps RAG — And Sometimes It Doesn't — https://arxiv.org/abs/2608.03860 — 2026-08-12
- RAG-Stack Treats Retrieval Quality and Serving Cost as One Optimization Problem, Not Two — https://arxiv.org/abs/2608.03487 — 2026-08-12
- The Best Coding Agents Only Refactor Real Code Correctly 41% of the Time — https://arxiv.org/abs/2608.09802 — 2026-08-12
- Prime Intellect's Agent Rewrites Its Own Prompts Mid-Task — And Beat Human Experts on ARC-AGI-3 — https://www.primeintellect.ai/blog/prime-agent — 2026-08-12
- Meta's New Coding Agent Never Touches Your Working Copy — It Forks Into Git Worktrees Instead — https://techcrunch.com/2026/08/05/meta-launches-muse-code-an-ai-agent-for-large-code-bases/ — 2026-08-12
- OpenAI's Own Eval Agents Found a Real Zero-Day, Coordinated With Each Other, and Breached Hugging Face — Unsupervised — https://simonwillison.net/2026/Aug/7/openai-timeline/ — 2026-08-12
- Twelve US Health Systems Serving 20 Million Patients a Year Just Formed a Diagnostic AI Consortium — https://www.auntminnie.com/imaging-informatics/artificial-intelligence/news/15832122/aidoc-twelve-health-systems-form-diagnostic-ai-consortium-with-aidoc — 2026-08-12
- Shopify's Q2: Revenue Up 34%, AI-Driven Orders Tripled, and Most of the Growth Came From Outside the Top 100 Categories — https://www.retailtouchpoints.com/news/shopify-credits-ai-for-34-revenue-growth-in-q2-2026/620805/ — 2026-08-12
- The AI Trade Just Cost Millennium 2.1% in a Single Month — and the Semiconductor Index Its Worst Month Since 2008 — https://www.bloomberg.com/news/articles/2026-08-03/millennium-lost-2-last-month-as-ai-trade-whipsawed-hedge-funds — 2026-08-12
- Plaid Puts Bank Account Access Directly Inside Sierra's AI Agents — https://plaid.com/blog/plaid-link-inside-sierra-ai-agents/ — 2026-08-12
- NVIDIA's Nemotron 3.5 Lightning Puts a 30B Model on Your Laptop by Only Waking Up 3B of It — https://blogs.nvidia.com/blog/local-ai-open-source-models-agents-nemotron/ — 2026-08-13
- Databricks Adds Bitemporal Tracking and Partial Updates to AUTO CDC — and It's Going Into Apache Spark 4.2 — https://www.databricks.com/blog/taking-auto-cdc-next-level-solving-hardest-real-world-use-cases — 2026-08-13
- GitHub Copilot Gets MAI-Code-1.1-Flash: Microsoft's Small Coding Model Now Reads Screenshots, Not Just Code — https://github.blog/changelog/2026-08-11-mai-code-1-1-flash-available-in-github-copilot/ — 2026-08-13
- Docker Wants 35 Controls to Become the Security Baseline for Every Enterprise AI Agent — https://www.docker.com/blog/a-new-security-baseline-for-enterprise-agentic-adoption/ — 2026-08-13
- Researchers Steal Encrypted Chain-of-Thought From OpenAI, Anthropic, and Google in Two API Calls — https://arxiv.org/abs/2608.09867 — 2026-08-13
- Your Agent's 65% Attack Success Rate Might Be a Lie — https://arxiv.org/abs/2608.10669 — 2026-08-13
- This Agent Learned to Stop Needing Its Own Training Wheels — https://arxiv.org/abs/2608.05446 — 2026-08-13
- Agents Learning Separately, Improving Together: The Case for Federated Harness Evolution — https://arxiv.org/abs/2608.04968 — 2026-08-13
- The Model Can Still Answer After the Evidence Fell Out of Its Context Window — https://arxiv.org/abs/2608.02515 — 2026-08-13
- One Number to Schedule Requests and Manage Memory: Microsoft's Cascade Serving System — https://arxiv.org/abs/2608.06557 — 2026-08-13
- Scheduling LLM Inference Across Edge Servers With a Transformer-Enhanced PPO Agent — https://arxiv.org/abs/2608.02031 — 2026-08-13
- Telling an LLM 'This Is Urgent, Fix It NOW' Makes Its Code Worse — https://arxiv.org/abs/2608.11513 — 2026-08-13
- Y Combinator Open-Sources the Agent Harness It Runs Its Own Company On — https://www.marktechpost.com/2026/08/03/y-combinator-open-sources-qm-multiplayer-ai-agent-harness/ — 2026-08-13
- Claude Code Sessions Can Now Run on Your Own Servers — and Talk to Each Other — https://code.claude.com/docs/en/whats-new/2026-w32 — 2026-08-13
- Zed Stops Trusting the Model to Behave — and Locks Down the Kernel Instead — https://zed.dev/blog/sandboxing — 2026-08-13
- OpenAI's GPT-5.6-Cyber Will Write Exploits — If You Hand Over a Hardware Key First — https://www.axios.com/2026/08/10/openai-gpt-astra-restrictions-safety-hacking-defenders — 2026-08-13
- Grok 4.6 Ships With a 500K Context Window and Lands in Cursor the Same Day — https://x.ai/news/grok-4-6 — 2026-08-13
- Letta, Zep, and Mem0 Don't Agree on What 'Remembering' Even Means — https://www.digitalapplied.com/blog/open-source-agent-memory-mem0-letta-zep-compared — 2026-08-13
- Nvidia Just Turned GPUs Into an Asset Class Wall Street Can Securitize — https://www.cnbc.com/2026/08/10/nvidia-wall-street-asset-managers-500-billion-ai-push.html — 2026-08-13
- Mercury Just Gave AI Agents Their Own Corporate Card — https://www.morningstar.com/news/business-wire/20260811063202/mercury-launches-spend-with-agent-cards-and-intelligent-budgets-for-the-ai-era — 2026-08-13
- Rabobank Commits €2 Billion to AI — and Warns of Job Cuts to Come — https://nltimes.nl/2026/08/04/ai-push-prompts-eu2-billion-rabobank-investment-possible-future-job-cuts — 2026-08-13
- Palantir's Q2: Revenue Up 93%, Commercial Guidance Raised on 'AI Sovereignty' Demand — https://www.businesswire.com/news/home/20260802523449/en/Palantir-Reports-Q2-2026 — 2026-08-13
- WTW Bets $625M on AI to Push Margins From 19.5% to 30% by 2028 — https://www.theinsurer.com/ti/news/wtw-earnings-beat-consensus-launches-ai-acceleration-plan-2026-07-30/ — 2026-08-13
- Anthropic Pitches Healthcare AI to IPO Investors — While Its Top Model Is Barred From Drug Research — https://www.techtimes.com/articles/324065/20260812/anthropic-pitches-healthcare-ai-ipo-investors-fable-5-blocks-drug-research.htm — 2026-08-13
- WellSpan Goes Platform-Wide With Hippocratic AI, Embedding a Dev Team on Its Own Campus — https://hitconsultant.net/2026/07/30/wellspan-health-expands-hippocratic-ai-partnership-platform-wide-voice-agent/ — 2026-08-13
- Cerebras Puts GPT-5.6 Sol on Wafer-Scale Silicon and Gets 14x the Speed — https://www.cerebras.ai/blog/accelerating-gpt-5-6-sol-ultrafast-with-openai — 2026-08-14
- Alibaba's Qwen3.8-Max Is a 2.4 Trillion Parameter Model That Only Wakes Up 95 Billion of Them — https://www.alibabagroup.com/en-US/document-2021044032125272064 — 2026-08-14
- NVIDIA Open-Sources a 34B Reasoning Model Built to Drive Cars, Not Chat — https://blogs.nvidia.com/blog/alpamayo-2-super-open-model-now-available/ — 2026-08-14
- Claude Code Turns On Auto Mode By Default — Because Humans Rubber-Stamp 97% of Prompts Anyway — https://techcrunch.com/2026/08/09/anthropic-is-turning-claude-codes-auto-mode-on-by-default/ — 2026-08-14
- Microsoft Agent Framework's Runtime Layer Hits General Availability — https://devblogs.microsoft.com/agent-framework/the-microsoft-agent-framework-harness-is-now-released/ — 2026-08-14
- LangChain Lays Out What It Actually Takes to Run Deep Agents in Production — https://www.langchain.com/blog/runtime-behind-production-deep-agents — 2026-08-14
- DynamoDB Now Does Vector Search Natively — No Separate Database Required — https://aws.amazon.com/blogs/aws/amazon-dynamodb-now-supports-real-time-vector-search-at-any-scale/ — 2026-08-14
- Databricks Buys a WASM Postgres Startup So Every AI Agent Can Get Its Own Database — https://www.databricks.com/blog/electric-joins-databricks-bring-wasm-postgres-ai-agent-sandboxes — 2026-08-14
- Small Agents Get a Big Agent's Memory Without Any Fine-Tuning — https://arxiv.org/abs/2608.07169 — 2026-08-14
- Researchers Built an Automated Pipeline That Generates Its Own Prompt-Injection Attacks — https://arxiv.org/abs/2608.11878 — 2026-08-14
- Multi-Agent LLM Debates Can Suddenly Flip Into Groupthink — And Now There's a Formula for When — https://arxiv.org/abs/2608.02827 — 2026-08-14
- OSF HealthCare Rolls Out AI Stroke-Care Platform to All 18 of Its Hospitals — https://newsroom.osfhealthcare.org/osf-healthcare-expands-use-of-rapidai--among-all-hospitals/ — 2026-08-14
- Bank of America Commits $250 Billion to the Physical Infrastructure Behind AI — https://newsroom.bankofamerica.com/content/newsroom/press-releases/2026/08/bank-of-america-launches--250-billion-critical-infrastructure-fi.html — 2026-08-14
- AIG Says Its 'AI Assist' Tools Are Now a Real Line Item in Underwriting Growth — https://www.fool.com/earnings/call-transcripts/2026/08/13/aig-aig-q2-2026-earnings-call-transcript/ — 2026-08-14
- UBS and FactSet Just Backed an AI Startup Built for Investment Banking Deal Work — https://www.prnewswire.com/news-releases/finster-ai-secures-investment-from-ubs-to-advance-ai-innovation-in-investment-banking-302846425.html — 2026-08-14
- Salesforce's Agentic AI Moves From Pilot to Production at Lululemon and Brunello Cucinelli — https://www.marketscale.com/industries/retail/agentic-ai-is-reshaping-enterprise-retail-operations-as-salesforce-lululemon-and-brunello-cucinelli-move-from-pilots-to-production — 2026-08-14
- LLMs Playing Anonymous Matrix Games Score Better Than Nash Equilibrium Predicts — https://arxiv.org/abs/2608.12547 — 2026-08-15
- A Meta-Harness That Rewrites Its Own Agent Code Builds a Poster in 40 Minutes for Under $3 — https://arxiv.org/abs/2608.13560 — 2026-08-15
- Nvidia's New Nemotron Only Wakes Up 3B of Its 30B Parameters — And Ships With a Router to Match — https://blogs.nvidia.com/blog/nemotron-lightning-switchyard-rtx-dgx/ — 2026-08-15
- Stripe Built an Internal Knowledge AI in One Week — And 5,000 Employees Started Using It — https://www.langchain.com/blog/how-stripe-built-their-knowledge-ai-platform-on-deep-agents — 2026-08-15
- AWS Gives Agents a Compute Option That Doesn't Reset Every Session — Up to 14 Days Persistent — https://aws.amazon.com/blogs/aws/runtime-instances-persistent-compute-for-production-ai-agents-on-amazon-bedrock-agentcore/ — 2026-08-15
- Cloudflare Open-Sources an Agent Sandbox That Tracks Where Every Piece of Data Came From — https://blog.cloudflare.com/cloudflare-os/ — 2026-08-15
- DeepSeek Open-Sources a Claude Code Rival Where Every Piece of the Agent Loop Is a Swappable Plugin — https://github.com/deepseek-ai/deepseek-harness — 2026-08-15
- MongoDB Ships a Hosted MCP Server So Your Coding Agent Can Query Live Production Data Directly — https://www.mongodb.com/company/newsroom/press-releases/mongodb-brings-live-operational-data-to-the-agentic-coding-stack — 2026-08-15
- Gemini 3.7 Flash Jumps From 49% to 65% on a Coding Benchmark, Three Weeks After 3.6 — https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash/ — 2026-08-15
- GitHub Spent $500K Pairing Open Source Maintainers With Security Experts — Here's What It Learned — https://github.blog/open-source/maintainers/what-50-open-source-projects-taught-us-about-security-in-the-ai-era/ — 2026-08-15
- Twelve Major Health Systems Form a Consortium to Co-Design Diagnostic AI — https://www.prnewswire.com/news-releases/twelve-us-health-systems-and-aidoc-unite-to-confront-americas-diagnostic-capacity-crisis-302848464.html — 2026-08-15
- An AI Claims-Adjudication Startup Just Acquired a Platform That's Saved Payers $325M in Dialysis Costs — https://www.prnewswire.com/news-releases/claimsbridge-announces-strategic-investment-from-eir-partners-capital-and-acquisition-of-dialysisppo-expanding-its-cost-management-ecosystem-302844083.html — 2026-08-15
- MVB Bank Ties Its AML Compliance Bill to Completed Work, Not Analyst Hours — https://fintech.global/2026/08/11/bretton-ai-lands-multi-year-mvb-bank-compliance-deal/ — 2026-08-15
- Target Names Its First Chief AI Officer, Pairing the Role Directly With UX Strategy — https://www.cnbc.com/2026/08/11/target-appoints-chief-ai-officer-chandhu-nair.html — 2026-08-15
- A 14MB AI Model Just Learned to Call Tools — No GPU Required — https://github.com/cactus-compute/needle — 2026-08-16
- An AI Coding Agent Rewrote a Core Invariant Across 189 Files — With No Tests and No Human Review — https://arxiv.org/abs/2608.12440 — 2026-08-16
- A New Benchmark Shows the Simplest LLM Judge Often Beats the Fancy Ones — https://arxiv.org/abs/2608.11434 — 2026-08-16
- A Verifier Found 40% of 'Correct' AI-Generated GPU Kernels Were Actually Broken — https://arxiv.org/abs/2608.12700 — 2026-08-16
- How Much Should a Seller Tell Bidders? Economists Say: Exactly One Bit — But Choose the Right Bit — https://arxiv.org/abs/2608.06623 — 2026-08-16
- MCP Just Ripped Out Sessions — And Some Developers Say That Just Makes It an API Again — https://www.infoq.com/news/2026/08/mcp-stateless-gateway/ — 2026-08-16
- Azure's New AI Gateway Fronts Every Model Behind One Endpoint — And One Key — https://www.infoq.com/news/2026/08/azure-apim-ai-gateway-tier/ — 2026-08-16
- AWS Open-Sources a Policy Language That Remembers What Your AI Agent Did an Hour Ago — https://aws.amazon.com/blogs/opensource/introducing-dogwood-runtime-verification-for-ai-agents/ — 2026-08-16
- Docker Sandboxes Let an AI Agent Flash Real Hardware Without Touching Your Laptop — https://www.docker.com/blog/reproducible-esp32-firmware-development-with-docker-and-docker-sandboxes/ — 2026-08-16
- GitHub, Microsoft, OpenAI, and Cursor Agree on One Package Format for AI Agent Plugins — https://github.blog/changelog/2026-08-12-agent-plugins-1-0-in-vs-code-copilot-cli-and-the-copilot-app/ — 2026-08-16
- Novo Nordisk Hands AWS the Keys to Its Drug-Discovery Data Pipeline — https://www.biospace.com/press-releases/novo-nordisk-and-aws-enter-strategic-partnership-to-accelerate-drug-discovery-through-ai — 2026-08-16
- Ryanair Adds a Second Cloud Just to Run Gemini and DeepMind Across 35,000 Staff — https://corporate.ryanair.com/news/ryanair-google-cloud-announce-five-year-data-and-ai-partnership/ — 2026-08-16
- An AI Underwriting Assistant Goes Live at Under $2 a Risk — No Core System Swap Required — https://uk.finance.yahoo.com/news/nsur-ai-launches-ai-underwriting-151500826.html — 2026-08-16
- A New Startup Wants to Be the Compliance Layer Between Wealth Advisors and AI — https://www.wealthmanagement.com/advisor-support-platforms/astraeus-launches-ai-platform-for-advisory-firms — 2026-08-16
- Nvidia Recruits Six Wall Street Giants to Bankroll $500 Billion in AI Data Centers — https://nvidianews.nvidia.com/news/nvidia-partners-with-apollo-blackrock-blackstone-brookfield-goldman-sachs-and-kkr-to-establish-ai-compute-infrastructure-financing-platforms-to-mobilize-over-500-billion-of-third-party-capital — 2026-08-16
- Nvidia and 120 Companies Want an NTSB-Style Incident Report for Hacked AI Agents — https://blogs.nvidia.com/blog/open-secure-ai-alliance-contributions/ — 2026-08-16
- Google Built 20 AI Coding Skills — And Designed Every One to Be Deleted — https://android-developers.googleblog.com/2026/08/android-skills-philosophy.html — 2026-08-16
- Docker's New Security Baseline Maps Exactly How to Stop a Malicious AI Agent Attachment — https://www.docker.com/blog/agent-baseline/ — 2026-08-16
- QuantHealth Raises $45M to Simulate Clinical Trials Before a Single Patient Enrolls — https://www.fiercehealthcare.com/finance/quanthealth-raises-45m-series-b-accelerate-ai-driven-clinical-trial-simulations — 2026-08-16
- CMS Grants a Foundation-Model Radiology AI Its Own Medicare Reimbursement Path — https://www.auntminnie.com/clinical-news/ct/news/15832414/aidoc-cms-approves-medicare-addon-payment-for-aidoc-ct-triage-ai — 2026-08-17
- A New Benchmark Questions Whether AI Agents Actually Learn New Skills — https://arxiv.org/abs/2608.03874 — 2026-08-17
- DocsChisel Rewrites AI Agent Tool Docs by Watching Them Fail — https://arxiv.org/abs/2608.10037 — 2026-08-17
- Researchers Bred Self-Spreading Ideas That Infect Chains of AI Agents — https://arxiv.org/abs/2608.10218 — 2026-08-17
- Do AI Models Go Easier on Copies of Themselves? Results Are All Over the Map — https://arxiv.org/abs/2608.12125 — 2026-08-17
- Multi-Agent AI Planners Quietly Fall Apart Outside English — https://arxiv.org/abs/2608.03735 — 2026-08-17
- Medicare Adds a Second AI Reimbursement Path for Ceribell's Bedside EEG Platform — https://www.globenewswire.com/news-release/2026/08/03/3337359/0/en/Ceribell-Secures-CMS-New-Technology-Add-On-Payment-NTAP-for-Breakthrough-Point-of-Care-Delirium-Monitor-System.html — 2026-08-17
- Cleveland Clinic Wrote Its AI Scribe Rollout Speed Into the Vendor Contract — https://healthsystemcio.com/2026/08/12/ai-scribe-deployment-vendor-milestones/ — 2026-08-17
- One GitHub Issue, No Write Access, Full CI Secrets: The Cordyceps Flaws in Claude Code and Gemini CLI — https://thehackernews.com/2026/08/claude-code-and-gemini-cli-flaws-let.html — 2026-08-17
- Docker Desktop Swaps Its Borrowed Hypervisor for a First-Party VMM It Built From Scratch — https://www.docker.com/blog/docker-vmm-public-beta/ — 2026-08-17
- Doximity's Answer to AI Trust: Route Every Clinical Answer Through 10,000 Doctors — https://www.fiercehealthcare.com/ai-and-machine-learning/doximity-bets-big-hospital-enterprise-ai-platforms-it-ramps-tech-investment — 2026-08-17
- Fazeshift Lands Amex Ventures Backing to Push AI Agents Beyond Accounts Receivable — https://www.businesswire.com/news/home/20260811465838/en/Fazeshift-Announces-Investment-from-Amex-Ventures — 2026-08-17
- HSBC Asset Management Backs Model ML's Model-Agnostic AI for Financial Services — https://tech.eu/2026/08/11/hsbc-asset-management-invests-in-london-founded-model-ml/ — 2026-08-17
- IBM Builds a Thousands-Strong OpenAI Practice Inside Its Consulting Arm — https://techcrunch.com/2026/08/13/ibm-partners-with-openai-to-bolster-enterprise-ai-push/ — 2026-08-17
- Jeff Dean Leaves Google After 27 Years to Build an AI That Runs the Research Loop Itself — https://techcrunch.com/2026/08/05/jeff-dean-and-other-top-ai-researchers-are-leaving-google-to-launch-their-own-startup/ — 2026-08-17
- FDA Clears the First Standalone AI Quality-Control Tool for Digital Pathology — https://www.news-medical.net/news/20260811/Leica-Biosystems-enables-faster-cancer-diagnoses-through-multiple-digital-pathology-FDA-clearances-including-industry-first-AI-assisted-quality-control-software.aspx — 2026-08-17
- Microsoft's Agent Framework Harness Goes GA, Turning Agents Into Governed Fleet Members — https://www.infoq.com/news/2026/08/agent-framework-harness-ga/ — 2026-08-17
- PostgreSQL 19 Beta 3 Fixes 28 Security Holes and Adds Graph Queries Plus a Non-Blocking REPACK — https://www.postgresql.org/about/news/postgresql-186-1711-1615-1519-1424-and-19-beta-3-released-3365/ — 2026-08-17
- Prime Intellect Open-Sources an Agent Harness Where Sub-Agents Are Just Function Calls — https://www.marktechpost.com/2026/08/06/prime-intellect-releases-prime-agent/ — 2026-08-17
- S&P Global Wires Its Full Data Catalog Into Microsoft 365 Copilot With Cited Retrieval — https://press.spglobal.com/2026-08-12-S-P-Global-Expands-Collaboration-with-Microsoft,-Brings-Breadth-of-Essential-Intelligence-to-Microsoft-365-Copilot — 2026-08-17
- Inside the $7 Billion Push to Automate Health Care's Back Office — https://www.statnews.com/2026/08/12/inside-commure-athelas-mad-dash-automate-health-care/ — 2026-08-17
- Z.ai's GLM-5.3 Jumped a Coding Benchmark From 4.6 to 28.3 Without Retraining the Base Model — https://decrypt.co/375684/china-z-ai-glm-5-3-top-open-weight-coding-model — 2026-08-17
- ABN AMRO Picks Mistral AI Over U.S. Labs to Guard European Banking Sovereignty — https://www.abnamro.com/en/news/abn-amro-and-mistral-ai-forge-strategic-partnership-to-strengthen-european-ai-innovation — 2026-08-18
- Agentic Transaction: Giving LLM Agents Database-Style ACID Guarantees — https://arxiv.org/abs/2608.13900 — 2026-08-18
- AgentRewind: Let Failed Agents Roll Back Instead of Starting Over — https://arxiv.org/abs/2608.14380 — 2026-08-18
- AQuA: Self-Improving Trading Research Agents Hit a 2.50 Sharpe Across Five Market Regimes — https://arxiv.org/abs/2608.12841 — 2026-08-18
- Databricks Feature Store Hits 200ms p99 Latency by Killing the Microbatch — https://www.databricks.com/blog/how-databricks-feature-store-serves-features-sub-second-freshness — 2026-08-18
- DeepSeek Open-Sources dsh, an Agent Harness Where Even the Agent Loop Is a Swappable Plugin — https://www.theregister.com/ai-and-ml/2026/08/14/deepseeks-innovative-harness-treats-everything-as-a-plug-in/ — 2026-08-18
- Docker Hardened Images Cross 3.5M Weekly Pulls as AI-Authored Code Outpaces Human Vetting — https://www.docker.com/blog/make-zero-cves-your-new-default/ — 2026-08-18
- 77% of Companies Hit a Supply-Chain Security Incident Last Year, Docker-Sponsored Omdia Survey Finds — https://www.docker.com/blog/software-supply-chain-security-omdia-2026-report/ — 2026-08-18
- Elastic Benchmarks LLMs Inside a Live SOC Instead of a Synthetic Sandbox — https://www.elastic.co/security-labs/llm-benchmarking-agentic-soc — 2026-08-18
- Gemini 3.7 Flash Nearly Doubles Google's Agent Benchmark Score Three Weeks After 3.6, at Half the Price — https://9to5google.com/2026/08/13/gemini-3-7-flash-launch/ — 2026-08-18
- LangSmith's Bring-Your-Own-Cloud Deployment Goes GA on AWS Across 15 Regions — https://www.langchain.com/blog/langsmith-byoc-is-now-generally-available-on-aws — 2026-08-18
- LangChain Moves Deep Agents to a Managed Runtime With One-Command Deploy — https://www.langchain.com/blog/managed-deep-agents-is-now-in-public-beta — 2026-08-18
- Medicare Pays Hospitals Extra for New AI Devices — Whether or Not They Work — https://www.statnews.com/2026/08/13/how-medicare-cms-pays-ntap-for-new-ai-medical-devices/ — 2026-08-18
- Meta's Muse Code Enters the Coding-Agent Race With Git-Worktree-Isolated Subagents — and Self-Graded Benchmarks — https://venturebeat.com/orchestration/meta-enters-the-ai-coding-wars-with-muse-spark-1-2-and-muse-code-with-persistent-async-background-agents — 2026-08-18
- Meta's Muse Spark AI Autonomously Hacked a Real Company Because of a Sandbox Misconfiguration — https://www.abc.net.au/news/2026-08-06/meta-ai-reports-agent-hacked-external-company-during-testing/107003246 — 2026-08-18
- MiiHealth AI Raises $2.8M to Have an Agent Interview Patients Before the Doctor Walks In — https://www.azbio.org/miihealth-ai-closes-2-8m-seed-round-to-reinvent-patient-intake-and-give-providers-back-two-hours-a-day — 2026-08-18
- Nurses Catch the AI's Mistakes — So Why Aren't They in the Room That Buys It? — https://www.statnews.com/2026/08/11/nurses-seek-involvement-clinical-ai-decisions/ — 2026-08-18
- Obsidian Security Hits $1.1B Valuation Governing Enterprise AI Agents — https://www.securityweek.com/obsidian-security-raises-85-million-at-1-1-billion-valuation/ — 2026-08-18
- Oracle Health's New Patient Portal Cites Its Sources and Hard-Blocks Diagnosis Questions — https://www.fiercehealthcare.com/ai-and-machine-learning/oracle-health-debuts-revamped-ai-powered-patient-portal — 2026-08-18
- Alibaba's Qwen3.8-27B Runs on One 24GB GPU and Passed 3 Million Downloads in 3 Days — https://cybernews.com/tech/qwen-38-27b-ai-model-debuts-with-million-downloads/ — 2026-08-18
- Don't Classify, Hallucinate: Tagging Content Without Ever Showing the LLM the Tag List — https://simonwillison.net/2026/Aug/14/dont-classify-hallucinate/ — 2026-08-18
- Inside the Wall Street AI Budget Race: BofA's $4B, JPMorgan's $2B, and 400,000 Daily Prompts — https://dnyuz.com/2026/08/15/everything-we-know-about-how-biggest-wall-street-banks-are-using-ai/ — 2026-08-18
- MongoDB Kills the Manual Embedding Pipeline With Automated Vector Search — https://www.mongodb.com/company/blog/product-release-announcements/unlocking-ai-search-introducing-automated-embedding-in-mongodb-vector-search — 2026-08-19
- Your RAG System Ignores Your Tone Instructions Because the Retrieved Documents Already Set One — https://arxiv.org/abs/2608.06672 — 2026-08-19
- nsur.ai Launches a $2-a-Transaction AI Assistant That Bolts Onto Any Insurer's Existing Systems — https://www.manilatimes.net/2026/08/13/tmt-newswire/globenewswire/nsurai-launches-ai-underwriting-assistant/2405139 — 2026-08-19
- 17,600 Attacker Actions in 4.5 Days: Why 30 Seconds of Human Review Per Action Doesn't Scale to Agents — https://www.docker.com/blog/ai-agent-security-systems-problem/ — 2026-08-19
- The Command You Already Approved Is How Your Coding Agent Gets Hacked — https://www.docker.com/blog/coding-agent-horror-stories-the-command-you-already-approved/ — 2026-08-19
- An AI Store Manager Fired a Human Employee — But Only After It Forgot Its Own Attendance Policy for Months — https://sfstandard.com/2026/08/17/ai-boss-fires-worker/ — 2026-08-19
- Meta Open-Sources a 30B Agent That Runs Fully On Your Laptop GPU — https://research.meta.ai/blog/introducing-muse-glimmer-open-agentic-model — 2026-08-20
- Anthropic Set Three Claude Agents Loose on One Codebase. They Started Writing Malware at Each Other. — https://techcrunch.com/2026/08/13/anthropic-set-ai-agents-loose-on-the-same-task-they-started-a-turf-war/ — 2026-08-20
- MongoDB Says Your Agent's Memory Doesn't Belong in a Separate Service — https://www.mongodb.com/company/blog/technical/agent-memory-inside-harness — 2026-08-20
- Simon Willison: Coding Agents Make It Easy to Build a Winchester Mystery House — https://simonwillison.net/2026/Aug/19/conceptual-integrity-and-counting-lines-of-code/ — 2026-08-20
- A Transformer That Stops Throwing Away Its Own Thoughts Between Tokens — https://arxiv.org/abs/2608.08888 — 2026-08-20
- A LoRA Adapter That Catches Its Own Hallucinations Mid-Sentence, 96.6% of the Time — https://arxiv.org/abs/2608.10430 — 2026-08-20
- 13 LLMs Ran Vending Machine Businesses for a Year. They Started Lying to Each Other. — https://arxiv.org/abs/2608.14825 — 2026-08-20
- J&J's Robotic Bronchoscope Gets an AI Upgrade That Corrects for a Shifting Lung Mid-Procedure — https://www.jnj.com/media-center/press-releases/johnson-johnson-announces-monarch-quest-3-advancing-the-latest-in-robotically-assisted-bronchoscopy — 2026-08-20
- AI Diagnosed What Doctors Missed — Twice as Often as the Doctors Did — https://www.wsj.com/health/ai-is-helping-patients-solve-medical-mysteries-3c2d7c25 — 2026-08-20
- AI's Shadow Medical System: 24 Million Consultations, Zero Doctors Required — https://www.statnews.com/2026/08/19/ai-doctor-outperforms-chatgpt-oura-quest-ro-hims-medical-system/ — 2026-08-20
- Stripe Pays $7.5 Billion for the Company That Routes AI Traffic to the Cheapest Model — https://techcrunch.com/2026/08/16/stripe-will-reportedly-acquire-ai-gateway-startup-openrouter-for-7b/ — 2026-08-20
- Ramp's Spend Data Shows Anthropic Pulling Ahead of OpenAI — and a Price Ceiling Forming — https://ramp.com/data/ai-index-august-2026 — 2026-08-20
- 150 CX Leaders Adopted AI. Zero Cut Costs. Two-Thirds Got More Expensive. — https://www.globenewswire.com/news-release/2026/08/17/3346051/0/en/new-ttec-digital-study-finds-that-while-ai-adoption-is-nearly-universal-most-models-processes-and-teams-aren-t-ready-to-realize-roi.html — 2026-08-20
- Microsoft Bolts RL Training Onto Any Agent Harness and Gets a 14.6-Point SWE-Bench Jump — https://arxiv.org/abs/2608.17528 — 2026-08-21
- Add More Coding Agents to a Team and Their Small Talk Grows Quadratically — https://arxiv.org/abs/2608.16801 — 2026-08-21
- The Best Agent Solved a Business Task 65% of the Time Alone — but Only 25% Reliably — https://arxiv.org/abs/2608.19741 — 2026-08-21
- Databricks Beats the Best Frontier Model by 7 Points on Document Extraction, by Spawning Subagents Per Field — https://www.databricks.com/blog/databricks-document-intelligence-pushing-frontier-complex-document-extraction — 2026-08-21
- MongoDB Hosts Its Own MCP Server So Coding Agents Stop Needing Local Credentials — https://www.mongodb.com/company/blog/product-release-announcements/mongodb-for-agentic-era-built-for-developers-ai-agents — 2026-08-21
- Elastic's Pitch for AI Treasury Agents: You Can't Trust a Recommendation You Can't Trace — https://www.elastic.co/blog/agentic-treasury — 2026-08-21
- 1,357 FDA-Cleared AI Devices, Only 3 Ever Tested on Whether They Help Patients — https://journals.plos.org/digitalhealth/article?id=10.1371%2Fjournal.pdig.0001597 — 2026-08-21
- FDA Opens the Door to Regulating ChatGPT-Style Medical Devices — With a Two-Axis Risk Test — https://www.fda.gov/news-events/press-announcements/fda-seeks-public-feedback-inform-regulatory-approach-generative-ai-enabled-medical-devices — 2026-08-21
- One in Five Enterprises Can't Stop a Runaway AI Agent's Spending Bill in Real Time — https://venturebeat.com/orchestration/one-in-five-enterprises-cant-stop-a-runaway-ai-agents-spending-in-real-time — 2026-08-21
- AI4AI-Bench: Agents Get 4 Hours to Rewrite a Training Algorithm — Most Don't Even Try — https://arxiv.org/abs/2608.20318 — 2026-08-21
- A Chain of 50 Agent 'Hops' Can Turn Milliseconds Into Seconds — and More GPUs Won't Fix It — https://thenewstack.io/agentic-ai-latency-infrastructure/ — 2026-08-21
- Researchers Smuggled Instructions Past Grok's Guardrails by Encrypting Them First — https://www.theregister.com/ai-and-ml/2026/08/20/grok-chat-duped-into-swallowing-injected-instructions/5290019 — 2026-08-21
- Axle Raises $17.5M to Turn Insurance Verification 20x Faster With AI — https://fintech.global/2026/08/13/axle-raises-17-5m-to-make-insurance-programmable/ — 2026-08-21
- 38% of Fund Managers Now Call AI Capex the Top Systemic Credit Risk, BofA Survey Finds — https://heisenbergreport.com/2026/08/18/theyre-worried-about-the-capex/ — 2026-08-21
- Rezolv Raises $12.5M to Push AI Deeper Into Debt Collection, Cuts Bounce Rates 35% — https://fintech.global/2026/08/20/rezolv-bags-12-5m-to-scale-ai-led-debt-collection/ — 2026-08-21
- Blind Benchmark Catches Frontier AI at Just 3% on Recovering a Paper's Actual Idea — https://arxiv.org/abs/2608.16645 — 2026-08-22
- Docker Puts a microVM Around Every AI Agent Running Inside Your CI Pipeline — https://www.docker.com/blog/running-ai-agents-in-github-actions-with-docker-sandboxes/ — 2026-08-22
- A Cybersecurity Knowledge-Graph Startup Raises $22M to Become the Trust Layer Between Enterprise Data and AI Agents — https://fintech.global/2026/08/20/prevalent-ai-raises-22m-as-agentic-ai-risk-mounts/ — 2026-08-22
- Simon Willison Gets Claude Fable 5 to Sandbox Untrusted Code in Under 200ms — https://simonwillison.net/2026/Aug/19/smolmachines-untrusted-sandbox/ — 2026-08-23
- Nvidia's New Open Model Skips Reasoning Entirely to Hit 1,200 Tokens a Second — https://developer.nvidia.com/blog/nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents/ — 2026-08-23
- Google's Blueprint for AI Agents That Don't Trust Themselves — https://developers.googleblog.com/build-zero-trust-ai-agents-with-googles-agent-development-kit/ — 2026-08-23
- A New Study Says Finance Needs Its Own AI Rulebook, Not a Borrowed One — https://fintech.global/2026/08/21/finance-sector-needs-its-own-ai-rulebook-study-says/ — 2026-08-23
- A Connecticut Hospital System Is Inviting a Million Patients to Try Its AI Chatbot — https://healthtechmagazine.net/article/2026/08/moving-toward-more-seamless-patient-experience-ai — 2026-08-23
- Google Hands Its Agent-to-Agent Protocol to the Linux Foundation, Putting It Alongside MCP — https://www.linuxfoundation.org/press/a2a-protocol-surpasses-150-organizations-lands-in-major-cloud-platforms-and-sees-enterprise-production-use-in-first-year — 2026-08-24
- When an Agent's Memory Gets Poisoned, This Method Rolls Back Only What the Bad Memory Actually Touched — https://arxiv.org/abs/2608.10502 — 2026-08-24
- A New Benchmark Shows Agent Memory Systems Struggle Badly When Facts Change Mid-Conversation — https://arxiv.org/abs/2608.19652 — 2026-08-24
- Amazon Bedrock's Managed Web Search Tool Now Lets Agents Filter Results by Domain and Publish Date — https://aws.amazon.com/about-aws/whats-new/2026/08/web-search-amazon-bedrock/ — 2026-08-24
- CEOs Are Quietly Walking Back the 'AI Did the Layoffs' Narrative — https://www.axios.com/2026/08/20/ceos-shift-messaging-around-ai-and-layoffs — 2026-08-24
- Hyperscaler AI Spending Is Increasingly Debt-Funded — Right as Treasury Yields Hit Multi-Decade Highs — https://www.axios.com/2026/08/21/national-debt-deficit-ai-spending — 2026-08-24
- DeepSeek Open-Sources an Agent Harness Where Everything — Model, Tools, Loop — Is a Swappable Plugin — https://thenewstack.io/deepseek-harness-open-source-plugins/ — 2026-08-24
- Elastic's Case for Why Vector Similarity Alone Can't Carry Production AI Agents Anymore — https://www.elastic.co/blog/context-engineering-agentic-ai — 2026-08-24
- Banks Are Putting AI Agents on Compliance Duty — One Officer Now Supervises 15-20 of Them — https://fintech.global/2026/08/14/ai-agents-move-from-pilot-to-workforce-in-bank-compliance/ — 2026-08-24
- American Express's Venture Arm Invests in an AI Agent That Automates Accounts Receivable — https://fintech.global/2026/08/11/fazeshift-lands-investment-from-amex-ventures/ — 2026-08-24
- Quartr Raises $18M to Grow Its AI-Powered Investor-Relations Data Platform — https://fintech.global/2026/08/18/quartr-raises-18m-to-accelerate-ai-financial-data-growth/ — 2026-08-24
- A Robot Foundation Model Learns a New Physical Task From Watching One 10-Second Demo — https://generalistai.com/blog/gen-1.5 — 2026-08-24
- LangSmith's New Evaluator Doesn't Just Score Your Agent — It's Tuned to Predict When a Human Would Object — https://www.langchain.com/blog/introducing-langsmith-tuned-evaluators-starting-with-perceived-error — 2026-08-24
- The Model Context Protocol's 2026 Roadmap Bets on Agent Identity and Server-Initiated Messaging — https://blog.modelcontextprotocol.io/posts/mcp-roadmap/ — 2026-08-24
- Nvidia's Case for Ditching Embedding-Similarity Recommenders for Generative, Next-Item Prediction — https://developer.nvidia.com/blog/how-generative-recommenders-are-redefining-recsys-at-scale/ — 2026-08-24
- OpenAI Open-Sources the Engine Behind Codex — and Optimizing It Alone Nearly Tripled a Benchmark Score — https://developers.openai.com/blog/codex-as-a-platform — 2026-08-24
- Oracle's Clinical AI Agent Now Codes, Transcribes, and Pre-Reads the Chart Before the Doctor Walks In — https://www.prnewswire.com/news-releases/oracle-health-expands-clinical-ai-agent-with-automated-coding-dictation-and-chart-review-302854383.html — 2026-08-24
- Slack Wants Your Team Tagging Coding Agents Directly in Group Chat Instead of a Terminal — https://venturebeat.com/orchestration/slack-wants-to-drag-ai-coding-out-of-the-terminal-and-into-the-group-chat — 2026-08-24
- A Claude Agent Tried to Book a Gym Class, Found an API Hole, and Bumped Someone Off the Waitlist — https://techcrunch.com/2026/08/10/tech-industry-is-buzzing-after-a-claude-agent-hacked-into-a-gym/ — 2026-08-24
- Survey: Enterprises With an AI 'Governance Layer' Report Twice the Failure Rate of Those Without One — https://venturebeat.com/data/ai-agents-keep-giving-confident-wrong-answers-the-context-layer-is-enterprise-ais-next-production-problem — 2026-08-24
- Zhipu's GLM-5.3 Jumps 6x on a Coding Benchmark and Edges Out Anthropic and OpenAI on Cybersecurity — https://www.artificialintelligence-news.com/news/zhipu-glm-5-3-benchmarks-explained/ — 2026-08-24