Crowded AI engineering conference hall, silhouetted speaker on stage, sea of attendees with lanyards under warm industri

Article · 6 min read

AI Engineer World's Fair 2026: labs sell services, agents ship, evals consolidate.

OpenAI's $4B services venture with Capgemini, Bain, and McKinsey landed last month, with Anthropic following soon after. The 6,000 engineers in Moscone West are about to find out whether their employer is now the frontier lab's customer or its competitor.

Context

Moscone West, June 29 through July 2: 6,000 engineers, 300 speakers, 29 tracks, and the audience tilts senior this year (a dedicated leadership track joined the lineup, the labs sent senior people, and VPs of AI were called out as a stated audience in the official program copy). The venue stepped up from the Marriott Marquis where the 2024 and 2025 editions ran, which tells you the conference graduated.

The year's headline

The 2026 headline is that the frontier labs went into services. OpenAI launched a $4B venture with Capgemini, Bain, and McKinsey last month (built on its Tomoro acquisition, per CRN's May reporting). Anthropic announced its own version within weeks. For the practitioners in Moscone West, the implication is structural: the consultancy that signs your paycheck and the lab whose model you fine-tune are now the same kind of business. Whoever ships production agents at a lower cost basis wins the next leg of enterprise spend. The labs are betting integrator markup is what they can compress.

Sessions worth showing up for

The full speaker list lands in the last two weeks before the show, and with 300 speakers across 29 tracks the names alone won't be the filter. Watch the slots:

The SWE Agents track keynote. Last year's SWE Agents track was the most signal-dense room of the fair (per the LangChain subreddit's field-notes thread, June 8 2025). This year's question is whether the Claude- and Codex-style compounding curve held through 2026 or hit a maintenance wall, where more time goes to fixing what the agent broke than to shipping. Ask the keynote: what percentage of internal commits at your company now originate from an agent run, and how has that number moved since January.

The MCP track, year two. Model Context Protocol went from announcement at the 2025 fair to working-standard within twelve months. The 2026 question is who runs the MCP servers at production scale and who pays for the infra. Ask: where does MCP server hosting land in your 2027 budget line.

Leadership-track sessions with 'production' in the title. The new leadership track exists because VPs of AI started showing up to a developer event. Sessions with 'production' or 'shipped' in the title (rather than 'roadmap' or 'evaluating') are where reorg politics get aired. Skip the panel titles that read like a McKinsey deck.

Whatever Anthropic and OpenAI present on services. Both companies just stood up direct-services arms (see year's headline above). Their senior people will be on stage to recruit, not to demo. Ask: where is the line between what you sell as a product and what you sell as a project, and how is that line moving each quarter.

The evals track. Consumer Reports' AskCR team flagged RecSys and Search+Retrieval as their highest-yield rooms in their 2025 recap (innovation.consumerreports.org). This year evals is the unsexy core question: how do you grade an agent that goes off-script. The good talks here are technical and boring on the outside, devastating on the inside.

Breakouts with signal density

The room with the most foot traffic is rarely the room with the most signal. Watch:

Enterprise case-study slots. When a non-tech enterprise puts an engineer on stage to describe what they actually shipped to actual customers (Consumer Reports' AskCR team did this in 2025 and wrote it up afterwards), the talk is more useful than every panel of frontier-lab founders combined.

Side-event evals workshops. Evals consolidation is the 2026 sub-headline. The named platforms (Braintrust, LangSmith, the half-dozen others) are running workshops outside the main schedule. The workshop you can't find on the official site is usually the one worth finding.

Open-source agent-framework BoFs. LangChain, LangGraph, CrewAI, Agno, Pydantic AI, plus whatever launched last week. The tension between these projects (who owns memory, who owns tool-calling, who pays for the orchestrator) plays out in the birds-of-a-feather sessions, not the main-stage panels.

International-startup booths. Mistral, Japan's SoftBank-Sony-Honda foundation-model consortium (per Tech Insider's coverage of the $3.3B NEDO-backed bet), smaller European labs. The conversation about sovereign infrastructure (Mistral announced manufacturing partnerships with Airbus, BMW, and ASML last month, per Cyber Magazine) is louder outside the US frame than the SF conversation lets on.

Anything labeled procurement, contracting, or 'production in regulated industries.' These are the rooms where the labs' new services arms will actually win or lose. The deal mechanics matter more than the model benchmarks now.

Companies to track at the booths

Five booths that justify a deliberate walk-through:

Composio. What they say: agent tool-calling infrastructure (per their LinkedIn announcement targeting AIE 2026 for market exposure, June 23). What they're actually selling: the bet that the agent stack settles on a small number of integration layers, and they get to be one of them. Ask: how many of your top 10 customers also have an in-house version of what you do.

Anthropic. What they say: Claude Opus 4.7 generally available with advanced software-engineering improvements (per the April 16 announcement). What they're actually selling: the services arm. The booth headcount this year will tilt heavier on solutions architects than on research engineers, which tells you the next 18 months of revenue mix.

NVIDIA. What they say: NemoClaw, autonomous engineering agents (announced at GTC Taipei, per the NVIDIA Blog). What they're actually selling: the bet that customers who built H100 clusters for training keep buying them for inference, plus a defensive moat against the Cadence-style domain-specific-agent threat (per Barron's Cadence stock coverage). The NemoClaw demos matter less than which integration partners they bring.

Tribal AI (or whoever is the metadata-native-agents seed-stage darling this quarter). What they say: $10M seed for metadata-native agents (SiliconANGLE, May). What they're actually selling: the thesis that the data layer eats the agent layer. Worth a ten-minute conversation if only to triangulate what enterprise buyers are asking for that the model providers can't answer.

The eval platforms (Braintrust, LangSmith, Patronus, Inspect AI, plus whoever launched this week). What they say: their grading harness is better than yours. What they're actually selling: lock-in for the layer that decides whether your agent ships. Pick two, ask each how they handle the cases their competitor can't, then ask the competitor the same question.

Conversation patterns

Three things that will be argued in the hallway:

  1. Whether the labs' direct services arms eat the consultancy work. If you're at Capgemini or Bain right now, the firm that signs your paycheck just took a check from the lab whose model your team integrates. The hallway question is whether the lab keeps you in the loop or eats your margin in the next renewal cycle.
  2. Whether the SWE-agent compounding curve held through 2026. The honest answer lives in the post-keynote Slack channels, not the main-stage Q&A. Watch for the engineer who says 'we rolled it back' more than the one who says 'we shipped it.'
  3. Open-source frontier versus closed-API frontier. Llama 4, DeepSeek, the Qwen catalog, plus Mistral's sovereign-stack pitch (per Cyber Magazine's coverage of the Airbus, BMW, ASML partnerships). The closed-API side has the better models for another six months. The open side has the better integration story for everyone who doesn't trust a US lab to host their data.

What nobody is saying out loud:

The inference compute supply at the scale every roadmap on the schedule assumes does not exist yet, and the data-center build-out fight just became a national-security story. HackerNoon's June 15 brief on the 'access-control era' and the OpenAI threat-activity report from this month (suspected China-linked clusters drafting content about US data centers themselves) sit at opposite ends of the same trend line. Nobody on stage will say 'our agent roadmap is a grid problem,' because the labs don't want to admit the constraint and the government's solutions aren't ready for public discussion. The constraint is real anyway. The talk that pretends it isn't has a 36-month half-life.

The follow-up coda

By the Tuesday after Moscone West, every booth conversation that mattered will be a name with no context attached. The LinkedIn-add reflex fires the moment the badge scan finishes, the actual follow-up never lands, and the warm window closes the day you sit back down at your desk.

Met turns the four-day deck of business cards and badge scans into a tracked follow-up queue inside the 72-hour window when each conversation is still fresh. iPhone only, free to install.

Download for iPhone

Read by AI engineers, founders, and VPs of AI heading to Moscone West who want the intel before they fly.

Get Met in your inbox

Field notes on conference networking, follow-up timing, and what we ship next. No spam, no AI hype.

No spam. Unsubscribe anytime. Replies go to support@sailquery.com.