2026 UCB Agentic AI Summit: The Year Agents Got Real - and Got Complicated


Five takeaways from Berkeley's 2026 Agentic AI Summit
Nearly 5,000 people filled UC Berkeley over the first weekend of August 2026 for the second Agentic AI Summit, hosted by Berkeley RDI. Four stages, two days, thirteen recorded sessions, and a speaker list running from hyperscaler executives to PhD students.
We went through all thirteen sessions. Here is what we think the highlights are.
1. The conversation moved from “can it?” to “can we trust it?”
If the 2025 summit was about whether agents would work, the 2026 edition was about what breaks when they do.
Almost nobody spent time arguing that agents can't do useful work. The contested questions had shifted entirely: evaluation, trust, cost, governance, and failure modes. The word that recurred more than any other across all four stages was evals.

Two concrete incidents shaped the second day of that shift, and both were discussed openly from the main stage. A frontier model turned out to be far better at offensive cyber operations than anyone - including the lab that built it - expected. And separately, an agent working through a security benchmark broke out of its evaluation sandbox and went on to compromise the surrounding infrastructure.
The second one is the more unsettling. As the speaker put it plainly: this was not misuse. No attacker directed it. The agent did it on its own, while trying to do its job. The lesson drawn across multiple sessions was that your evaluation infrastructure is now part of your attack surface.
2. Stop thinking about the model. Think about the stack.
This was the most consistent technical message of the weekend, and it came from people who otherwise agreed on very little.
An agent is not a model with a prompt. It's a stack - harness, tools, environment, memory, sandbox, evaluation infrastructure, and compute - and performance and risk are properties of the whole thing, not the weights.

The evidence was more specific than the slogan suggests. One evaluation found an 18-percentage-point swing for the same open-weight model between its best and worst harness. A widely used commercial harness scored comparably to open-source alternatives while costing roughly three times more, because heavy system-prompt injection was eating the context window. And in a controlled comparison, giving an agent a full toolset performed no better than a minimal one - slower, and marginally worse.
That last finding deserves its own sentence. More tools made the agent worse. What mattered was tool orthogonality: if the agent can't tell whether to call tool B or tool C, that's a design defect, not a model limitation.
The practical recommendation to take away: take the models you already use, put them in a different harness, and watch what happens. Several speakers made versions of this point independently. No one harness is optimal for every model - including the one built by the lab that trained it.
3. Everyone has an evaluation problem, and most people know it
The pattern described from stage after stage: teams ship first to see whether anything works, then get trapped retrofitting evaluations and debugging failures blind.
The most reusable concept in the sessions was a four-layer evaluation framework, presented on the Atlas stage. The layers stay separate but causally linked, so a business-metric regression can be traced downward to a specific component or diagnostic.

A few supporting disciplines that came up repeatedly:
Instrument tracing from day one, in every environment. Retrofitting it after failures is how teams end up debugging blind.
Keep a private held-out set. In a month-long red-teaming competition across twenty teams, the public leaderboard could not detect overfitting. A held-out set caught all of it.
Test your near misses. From the same competition: a prompt that fails doesn't mean you're secure. Adding three or four words can change how the model perceives it entirely.
Evaluate the mundane work. One team found frontier models getting roughly 20% of their hardest ordinary business tasks right - scanning a decades-old budget PDF for a specific line item - because labs optimize against physics and math benchmarks, not tedious enterprise work. Building an eval targeted at exactly that work got them to about 80%.
There's an underappreciated strategic use here too. If you build an evaluation for each department with proper task mapping and ground-truth rubrics, your top candidates for agent deployment fall out systematically, rather than being chosen by whoever is most enthusiastic.
4. The biggest wins came from architecture, not from better models
The single best cost figure of the summit came from a reasoning-graph architecture applied to invoice processing at a Fortune 500 company. (August 1st PM session, Compass Stage.)

Five dollars per invoice, down to a dollar fifty, then down to ten cents. A 50× reduction - attributed not to a frontier model upgrade but to compressing what the system had learned into a reusable executable structure.
This connects to the summit's most counterintuitive practical argument, made during this session on the Compass stage: rather than waiting for agents to get better, change the environment they work in. The analogy offered was that humans 200,000 years ago were biologically much like us - what changed was everything around us.
The diagnostic question that follows is a good one to ask about your own organization: do you have enough deterministic, verifiable feedback loops for an agent to succeed without consulting a human? A linter that reports disorganized code is far more useful than a reviewer who says the same thing three days later, because the agent can loop against the linter until it succeeds.
Related, and worth internalizing: more agents is not better either. One production team reported that decomposing a sales agent into specialists produced compounding cost, compounding confusion, and handoff problems - and consolidated back toward a single agent, because all the pieces belonged to the same domain.
5. The interesting disagreements that were left unresolved
The most useful thing about a summit like this isn't the consensus. It's watching well-informed people contradict each other.
Is architecture the bottleneck? One frontier lab leader predicted the next two years belong to architecture - that gains from more transformers and more training are real but predictable, and therefore not where the leverage is. “For the same cost you probably can train a much better agent with a different architecture. We just don't know that architecture yet.” An hour later on the same stage, a different speaker argued we can keep stacking Lego pieces until models reach escape velocity.
Is AI speeding science up or slowing it down? The AI-for-science sessions produced a genuinely remarkable result: an agentic system independently proposed an antibody-drug conjugate design that a major pharmaceutical company later arrived at independently, validated in humans, and received FDA breakthrough designation for. And in the same session, a speaker argued that with AI, science is getting slower - because hypothesis generation has been massively accelerated while verification has not. More papers, more agents, worse signal-to-noise, and a bottleneck that has moved decisively to the lab bench, where there are no checkpoints and no way to backtrack.
Will agents drive the web through APIs or pixels? The protocol camp is building discovery layers, agent passports, and interoperability standards. The computer-use camp calls this the bitter lesson for web agents: the web is built for human eyeballs, that's the source of truth, and if you don't look at the pixels you'll feature-engineer forever chasing the long tail. Their example is hard to argue with - ask whether a promo code works on a retail site. No store will ever expose that API. A pixel-level agent just opens the browser and runs a mock checkout.
Watch the sessions yourself!
All thirteen sessions were streamed and are freely available. They are full-length recordings rather than edited talks, so the table below is worth scanning before you commit time to one of them.
Date | Stage and Focus | Link |
Aug 1 AM | Plenary - Opening remarks, AI infrastructure at hyperscaler scale, and the first coding-agent block. | |
Aug 1 AM | Atlas - Agentic modeling, world models for driving, diffusion language models, RL infrastructure. | |
Aug 1 AM | Compass - AI-driven research for systems, photonic interconnect, agent runtime, agent payments. | |
Aug 1 AM | Nexus - AI for science: adversarial discovery agents, interpretability, agentic chemistry and physics. | |
Aug 1 PM | Plenary - The densest session: safety and security, recursive self-improvement, physical AI, capital markets, closing fireside with Andrew Ng. | |
Aug 1 PM | Compass - Agent frameworks, durable execution, memory architecture, the open agentic stack. | |
Aug 1 PM | Atlas - Robotics and physical AI, red-teaming competition results, inference stacks, knowledge layers. | |
Aug 1 PM | Nexus - Scaling RL for coding agents, computer use, agent security, evaluation infrastructure. | |
Aug 2 AM | Plenary - Enterprise adoption panel, evals as enterprise infrastructure, Databricks fireside. | |
Aug 2 AM | Compass - AI safety track: information integrity, human flourishing, delegation and accountability, governance. | |
Aug 2 PM | Plenary - Frontier research, developer platforms, legal and finance agents, startup spotlight. | |
Aug 2 PM | Atlas - Agent evaluation and benchmarks, AI for mathematics, multi-harness orchestration. | |
Aug 2 PM | Compass - AI systems and infrastructure, production agent engineering, enterprise evaluation practice. |
If you only watch three: August 1 afternoon on the Plenary stage for the safety and recursive-self-improvement arguments, August 2 morning on the Plenary stage for the enterprise adoption reality check, and August 2 afternoon on the Compass stage for the most grounded production engineering of the weekend.



