← 返回本期Every 联合创始人兼 CEO Dan Shipper 以他所谓的「技术革命时期照常营业的极度无效」开篇:Fable 5.1、Astra 6 等前沿模型不断重置可能性边界,产品负责人既无法忽视,也不能任其扰乱路线图。他的诊断是,探索(发散、重演示、大部分产出会被丢弃)与执行(收敛、围绕既定路线图交付)是两种对立的工作方式,因此组织应通过实验室团队实现关注点分离:实验室在新模型发布时探索能力上限,预期丢弃约 90% 的产出;产品团队采纳其中约 10%,并保持上线产品的连贯性——AI 让一人实验室也成为可能。他给出具体做法:以一至两人的 two-slice 团队取代传统 two-pizza 团队,让海盗式的人(快速、凌乱的价值猎手)与架构师(把可行原型塑造成可扩展系统)配对;通过自己先用获得最紧的反馈回路;并行运行相互竞争的实验;把被丢弃的实验转化为对外内容或早期用户计划,使整体 ROI 为正。随后由研究管线把赢家从内部使用推向早期客户、再进入正式产品,晋升判据明确:是否有持续使用与回流、是否比现有方案好 10 倍、能否在规模化成本内提供服务。他用 Every 的 Kate Bench(自动化主编的文稿润色,使其此类编辑工作逐月减少 12%)以及 Anthropic Labs(Claude Code、MCP、Skills)与 OpenAI(Codex 从主线之外的小团队成长为 ChatGPT 这个 8 亿日活应用的基石)佐证这一循环正在真实运转。

视频 · Lenny's Podcast

在移动的 AI 前沿上做产品:用实验室机制驾驭技术变革 | Every CEO Dan Shipper

原题:How to build products on a moving frontier | Dan Shipper (Every)

Lenny's Podcast约 13 分钟
内容摘要Every CEO Dan Shipper 指出,探索 AI 前沿与执行路线图是两种根本对立的工作方式,产品组织应设立由海盗与架构师组成的一至两人实验室团队,通过并行实验与内部自用,把真正有效的约 10% 经研究管线推入主线产品。

Brief Description

Dan Shipper, co-founder and CEO of Every, explains how to build products when the AI frontier keeps shifting beneath your feet — what he calls "the unreasonable ineffectiveness of business as usual during technology revolutions." His argument: executing a roadmap and exploring the frontier are fundamentally opposed ways of working, so product organizations should add research-lab elements — tiny "two-slice" teams of pirates and architects running parallel experiments and dogfooding what they build. He walks through why labs matter, how to run one, and how to push winners through a research pipeline into the main product, illustrated by Every's effort to automate its editor-in-chief's copy edits and by Anthropic Labs (Claude Code) and OpenAI (Codex).

Table of Contents

  • The World Just Changed — Now What?
  • The Core Problem: Exploring vs. Executing
  • Why Make a Lab?
  • How to Run a Lab
  • The Research Pipeline
  • Case Study: Automating Every's Copy Editing
  • Keeping the Pipeline Moving
  • Merging Winners into the Main Product
  • Closing: Welcoming the Next Model Drop

The World Just Changed — Now What?

I want to talk about how to build products on a moving frontier — or, as I like to call it, the unreasonable ineffectiveness of business as usual during technology revolutions, and what to do about it.

The world just changed. The world just changed again. Fable 5.1 launched last week. Astra 6 launched last week. And these are seriously impressive models. I made a 3D historical reproduction of the Battle of Waterloo that's historically accurate with a single prompt — it just churned for four hours, and then I got that. That's crazy.

Someone else on my team made an agent simulation: he fed it a scientific paper about how agents interact and get simulated together, and then had a thousand agents all running on his computer and visualized, after just a couple of prompts. It's also starting to do our video editing. We do a lot of videos, and Astra is actually good enough to go into Premiere and do a bunch of editing. And fun fact: all of the animations in this deck, and a lot of the deck itself, were made by Astra — I didn't touch the animate tab at all.

So this is crazy. What should you do? If you're a product leader with a product team, do you keep your head down and focus? That might work for a little while. Do you ask your customers what they want? Maybe, but most of your customers are not familiar with this stuff — they're actually looking to you to tell them what to do. Do you turn your B2B SaaS app into a 3D multiplayer strategy game? Maybe. Maybe. But it's hard. This is a big problem that everyone in this room faces. I face it.

The first thing to know — rule number one — is never make any major life decisions within 30 days of a meditation retreat, a psychedelic experience, or your first encounter with a frontier model. All right, we're settled. That's rule number one.

The Core Problem: Exploring vs. Executing

Let's get into the problem. Why is this so hard? If you're running a product team, it feels like it's everyone's job to execute the roadmap at a high level and to stay at the frontier at the same time. That's really hard, because these are very opposed ways of working.

If you want to explore the frontier, you're exploring — that's divergent. You're doing lots of different things, you're probably going to throw a lot of it out, you're trying all the new models, you're doing demos. Exploiting is very different. You're converging, you're focusing, you're saying no to things, you're executing against a roadmap you've planned out. Doing both of those at once means you're pulling in opposite directions.

So what should we do? My contention is that you should run your product org like a research lab, or add some research-lab elements to it. We're going to go through why you should make a lab, how to run a lab, and how to bring what works from the lab back into your main product — using examples from what we do at Every and from some of the best companies in AI.

Why Make a Lab?

Right now, everyone in your org is doing both execution and exploration. But there are some people in your org — I guarantee it — who are doing way more exploring than others. It's really important to identify them, because they are your early adopters. We are your early adopters: we love new things, and we're already running Astra and Fable on the weekends for personal projects, or to explore things we might bring to work. Because of that, we'll have a really good sense of what your product might become, because we're already living in the future. The problem is that we can be a massive distraction.

So the key question, if you're running a product org, is: how do I harness the early adopters on my team — maybe even some of my customers who are early adopters — without distracting everybody else?

The solution we've found at Every, and that a lot of really great companies are starting to adopt, is to make a labs team. With a labs team you've separated concerns. Some people are in charge of improving and scaling what already works — they're on the product team. Some people are in charge of exploring what comes next: exploring the frontier, doing demos, trying new models. That lets you get the best of both worlds.

And if you're looking at this thinking, "That's really expensive, and I don't have the resources for it" — what's amazing about AI is that it allows you to have a labs team of one. You can have one person on your team whose job it is, right when a new model comes out, to go explore it and come back and tell you what they learned. It doesn't require a lot of resources, because everyone on your team now has so many superpowers — they can just have Astra and Fable do a bunch of stuff — that it doesn't take a huge investment.

To pull the difference apart: on a labs team, your job is to explore capabilities — especially when new models drop — to see what's now possible, and to run a ton of experiments in parallel. Crucially, the expectations are very different. On a labs team you expect to dispose of about 90% of what you make: you try it and you throw it away. A product team is very different: your job is to improve and scale the product, deliver for existing customers, and make sure they feel like they understand the product — it's coherent, and you're not just throwing a bunch of garbage at them. A product team should expect to adopt about 10% of what the labs team makes or tries, and that's how the two start to work together.

There are really good examples of this starting to work in AI. We've had labs teams for a long time, but it's starting to work in a way it never has before because of this technology. A great example is Anthropic Labs. Claude Code — one of the biggest, most successful productivity products of all time — came out of Anthropic Labs, a separate little group in charge of doing lots of experiments. MCPs, Skills, Claude Design — all of these came from a small group of people inside Anthropic experimenting, and then Anthropic investing more and more into the winners, the ones that worked. And there are a thousand experiments you've never seen.

How to Run a Lab

Say you're convinced this structure gets the best of both worlds — it unleashes your early adopters without distracting the people on the product team who need to deliver on the roadmap. How do you run a lab well?

Use very, very small teams. For a long time the standard for team size was the two-pizza team — eight to ten people, per Jeff Bezos. In the AI age, I call it a two-slice team: one or two people max. You can get so far so fast with one or two people that anything more brings coordination overhead and differing visions, and it just doesn't work.

The composition that works for us — and that I think is starting to work for others — I call pirates and architects. One person on the team is the pirate: a slop cannon absolutely obsessed with finding value. That's me — I'm out there doing lots and lots of stuff, throwing it away, all of it messy, but I'm going to find something really interesting. Architects are people who look at a messy system — maybe something I built that's a complete mess — and help shape it into something valuable, beautiful, and extensible. Pairing those types together is very, very powerful in AI.

Dogfood. The biggest thing that matters on a labs team is making the feedback loop as tight as possible between making something and knowing if it's good. The tightest feedback loop is making something for yourself. If you can do that, do it. If you can't, it's really important to get a couple of early, early-adopter-type customers to help you test with a tight feedback loop so you can iterate really fast. And use what you're building: ideally every experiment should be used for actual work you actually have, so you can tell whether it's useful or just new — that's a big distinction you need to draw.

Build many experiments in parallel, even if those experiments are trying to do the same thing. Try competing approaches to the problem. It might look like a mess, like a lack of coherence, but everybody will have a slightly different take on how to solve a particular problem. When capabilities move, the frontier becomes very unknown, and having people try the same thing from different perspectives helps you map that frontier and decide what's valuable and what's not.

Make your experiments net-positive in terms of ROI. We're going to throw out 90% of what we make — so how do we make even that 90%, the part that never makes it into the product, actually useful? One thing we've done a lot at Every that's worked really well is turning experiments into external content. People love to see everything we tried, what worked, and what didn't — and that brings customers to read our stuff and use our products. If that's possible, I highly recommend it. Another option is to use experiments to feed an early-adopter program: a lot of your customers want to be closer, want to be in the fold, and this can be a value-add that lets them work with you more closely. And at base, the labs team should be sharing what they learn about capabilities and what's now possible with the product team, so the product team can take that into account as they build — without getting distracted by having to explore the frontier themselves.

The Research Pipeline

So you've made a lab and started to figure out the best practices. How do you get what you're making from the lab into your product? The thing you need is a research pipeline: ideas start on the left as lab-only, with lots of little experiments, and you move them progressively from left to right. At some point you have a handoff with your product team, where they start to get incorporated into the product.

Here's what it looks like. The labs team runs lots and lots of experiments. Most of them are not good, but a few might work, and those become things you start to test in real work. One thing we do all the time internally is just let other people on the team adopt them and see if they like it. There are different configurations that work depending on what kind of customers you serve, but the first step for us is always: do other people on the team start to use it? If so, it's often ready to put in front of early customers, and ready for the product team to really take a look at — to figure out where it fits in our roadmap and how we might use it. Even there, not everything makes it. But at some point, the thing you're building will be ready to scale and release to customers.

Case Study: Automating Every's Copy Editing

Here's a concrete example. Kate is our editor-in-chief, who I've been trying to automate — very lovingly — for the last three years. Kate is fantastic at a lot of things, but one thing she's really good at is taste for copy edits. As the team has grown — we're about 30 people now — she can't do all the copy edits herself. We could hire someone, but it's really hard to find someone with her level of taste and train them. So what I've been trying to figure out is how to make AI help her, so she can expand her impact in the org without spending additional time doing copy edits late at night.

Kate often wakes up to messages from me like: "Not important, but I downloaded every single one of your copy edits over the last three years and had Fable try to do a copy edit on the latest piece." I've literally been doing stuff like this for a couple of years, and it just started to work — the capabilities are finally good enough that it's ready for Kate to actually use. We call it Kate Bench: we'd done a bunch of experiments, one was starting to work, and now we pushed it into the org to get some internal use.

We have an Every agent — a single agent for your entire company that helps your company adopt all of its AI workflows, at agents.every.to — and we use it internally. Now Kate can say, "At Every, do a Kate pass of this draft," and it will literally file suggested changes like she would, based on her historical edits, and improve over time. I built this in my spare time and it just sort of works. Now we're starting to see that there's value for the organization, because she's adopting it and it's spreading to other people — that's a really good sign.

When we get to this internal-use stage, that's the time to bring in an architect — I have one in particular on my team, Yannik, who's fantastic. Yannik takes the messy thing that kind of works and makes it really great. Now we have a whole dashboard: for each document, how many suggestions were accepted, and how much work remains for Kate after the agent does its pass — it checks the edits she makes after the agent has gone in. And you can see we're improving: she does 12% less work on these types of edits than she did the month before. His job is to make it a real system we can improve and compound on an ongoing basis. Now it's starting to work well enough that we're thinking about how to put it in front of early customers — and whether this process works for things beyond copy editing. I think it does. We've taken it from "only an experiment in the lab" to something we're ready to give to customers.

Keeping the Pipeline Moving

A few best practices for continually pushing things through your lab pipeline.

Review the pipeline regularly. I have a tracker in Notion, and every week at our all-hands, everyone talks about what's going on in the pipeline. That's crucially important. You want the rest of the product team — the people who aren't playing around with everything — to know what we think is interesting and what's starting to move up the pipeline, so they can start thinking about the implications for the product if it actually proves out, instead of having to scramble to run their own experiments. It keeps the lab and the product orgs in sync.

Define clear decision criteria for moving an idea through the pipeline, and really make it an event when you do. Some of our criteria: the big one is, are people using it, and are they coming back? What you're trying to do is use internal use as a proxy for value. The second big question is, is it 10x better than what currently exists? That's a really interesting filter, because what is new feels very exciting — but a month later, is it actually any better? It's hard to know unless you use time and usage over time as a filter. And also: is it affordable? Maybe it works now, but could we actually serve it at scale for our customers? That's another really big question. Having clear criteria lets you make sure you're being rigorous about how you do this.

Merging Winners into the Main Product

The ultimate goal is to merge the winners into the main product. This full cycle — labs to merging winners — is something that's starting to happen really, really well. We've had labs for a long time, and the idea of self-disruption — building something that then disrupts your own product — has been around for a while. But it's actually starting to work now in AI, because you can move so fast.

A really good example is Codex. Codex was built by a small team working outside of the main app. There were a bunch of different teams inside OpenAI working on the future of coding at the same time as Codex was, doing iterations on different form factors for what the future of coding would look like: Is it inside an IDE? Is it a CLI? They launched the desktop app, which I think is fantastic, in February 2026, and it grew so fast that they eventually just merged it into ChatGPT — it became the foundation of ChatGPT. The small little Codex team then took over this 800-million-daily-active-user app, because they were so successful with this way of experimenting.

Closing: Welcoming the Next Model Drop

I think that's the promise of doing something like this: you can actually build the next version of your product while scaling the one you already have — and you can do it without going crazy. The way you know it's working is that you will welcome the moments when new models drop and be excited about them, instead of dreading them.

And that's my talk. I'm Dan Shipper, co-founder and CEO of Every. You should check out Every — the only subscription you need to stay at the edge of AI. Thank you very much.