All posts
Blog

Can AI Predict Startup Success? What Two Months of a Live Forecasting Arena Taught Us

Ahead of launching VC Arena and VCRouter, we ran a two-month experiment: GPT, Claude, and Gemini predicting live startup outcomes, scored against what actually happened. Four early signals came out of the data, and each one shaped what we are building as the next version of VCBench.

Arenas rank what people like. Ours scores what comes true

If you follow AI, you have seen the arenas. Arena AI and LMArena rank LLMs through millions of human votes on anonymous answers. Design Arena does the same for AI-generated designs. These leaderboards work because for chat and design, taste is the right judge: if people like the answer, the answer is good.

Venture needs one more kind of arena. In investing, a brilliant-sounding prediction that turns out wrong is worth nothing, and no preference vote can catch that. The judge that matters is the future.

Preference arenas ask which answer people like. Ours asks which predictions come true.

We are building that arena, and a model router on top of it: VCRouter, releasing soon. Before building, we at Vela, a tier-1 quant VC, ran a two-month design study to learn what such an arena must measure. This post shares the setup, the early numbers, and where they point.

The experiment: two months, 858 predictions, graded by reality

We gave GPT, Claude, and Gemini live questions about the startup world, with answers nobody could know yet. Which of today's Product Hunt launches will finish first? Will this new Show HN project take off within 24 hours? Each model had to state a probability, not just a pick, because a forecast you cannot score for calibration is just an opinion. Then reality answered, and we graded every prediction against it.

78questionsLive questions resolved by real-world outcomes, June to August 2026
858predictionsStated-probability forecasts from GPT, Claude, and Gemini
~1¢eachProvider-billed cost per prediction. The measured subset cost $0.86 in total

Four early signals came out of the data. Each one shaped what we are building.

Signal 1: models already read the market extremely well

On the Product Hunt task, guessing at random wins about 1 time in 6. The best models won 9 times in 10.

Where does that skill come from? Mostly from reading. The product already in first place at question time went on to win 8 times in 10, and the models that leaned on that signal hardest scored highest. Today's frontier models are exceptional at scanning a noisy page, finding the strongest signal, and acting on it. That is a production ready capability right now: an analyst that instantly and reliably tells you who is winning in any market.

It also taught us the first design rule of a venture arena: separate reading the present from predicting the future, and score them as different skills. A benchmark that mixes them will overstate one and hide the other.

Signal 2: the gap is calibration, not intelligence

The Show HN task was a genuine forecast: the post was still climbing, and about 1 in 5 of these projects ended up succeeding. Think of a weather forecaster. If she says “70% chance of rain” across a hundred days, it should rain on about 70 of them. That property is called calibration. Here is how the models' numbers held up:

When the models saidIt happenedVerdict
“20 to 40% likely”15% of the timeOn target
“40 to 60% likely”50% of the timeOn target
“60 to 80% likely”14% of the timeToo optimistic
“80 to 100% likely”27% of the timeToo optimistic

The pattern is consistent: the models lean optimistic. They predicted success on half the questions when only 1 in 5 succeeded, and “about 70% likely” came true 1 time in 7.

Underneath the optimism, the promising part: on the models' most bullish calls, the hit rate ran at 1.9 times the rate of chance. Early and small-sample, stated as such. But it points at a clear thesis: the raw ranking ability is there, and what needs building on top is a calibration layer, trained and validated on resolved outcomes.

The gap between today's models and useful venture forecasting is not intelligence. It is calibration, and calibration is buildable.

Signal 3: the crowd of models beat every single model

The single best forecaster in our system was not a model. It was the average of all of them. The blended prediction ranked outcomes better than any individual model, including the best one.

Decades of research on human forecasters found the same thing: crowds beat champions. For anyone using AI in venture, the practical rule is simple. Do not ask which model is best; use several and blend them. The blend is nearly free, and it is more stable than any champion picked from a small sample.

This signal is the founding insight behind VCRouter: routing and blending across models, steered by live arena standings, instead of betting on one.

Signal 4: the economics already work

The predictions with exact provider-reported bills cost $0.86 in total, about a penny each or less. The most expensive model cost three times the cheapest with no measurable accuracy difference, and the cheapest model was also the fastest.

One practical footnote: web search did not help. Paired on the same questions, accuracy was identical to the percentage point, at a multiple of the cost:

ModelAccuracyCostSpeed
GPTUnchanged32x more4x slower
ClaudeUnchanged24x more3x slower
GeminiUnchanged4x more3.5x slower

For a question about the future, the answer is not on the web yet. So for forecasting tasks we run search off, which makes large-scale prediction even cheaper. Running thousands of scored predictions a month costs less than a lunch. The economics of this field are solved; what it needs is resolved outcomes at scale.

Where this is going: VC Arena and VCRouter

One more number shaped the roadmap. A statistically definitive verdict on forecasting skill takes roughly 800 resolved predictions; our pilot collected 85 on the hard task. So the next step is scale, and that is what we are building: VC Arena, the next version of VCBench (research summary). VCBench tested models on 9,000 historical founder profiles; the arena moves from frozen history, which can be memorized, to the live future, which cannot.

Here is a preview of what excites us about it:

  • Every venture task, one leaderboard. Sourcing (company and founder discovery), analysis (market maps, investment memos, founder evaluation), forecasting (seed to Series A within 18 months, GitHub breakouts, Product Hunt winners, M&A), and value creation (launch optimization, fundraise readiness). One standings table, updated as reality resolves.
  • Two kinds of scoring, honestly separated. Forecasts are graded by outcomes. Judgment work, like market maps, is judged in blind battles: two models, same brief, names hidden, voted on by verified investors and operators. Pick the map you would take to Monday partner meeting.
  • Questions models cannot memorize. Forecasting tasks resolve after each model's training cutoff, so the only way to score well is to actually predict.
  • The human crowd as a competitor. On every forecast, the aggregated human prediction runs as its own entrant. Man versus machine, on the record, every day.
  • VCRouter, releasing soon. One API call sends any venture task to whichever model currently leads it in the arena. As the leaderboard moves, your routing moves with it.

That is the quant VC way of working: publish the methods, score in public, and let reality keep the leaderboard honest. For the same discipline applied to human investors, see our analysis of how tier-1 firms pick founders.

Methodology and data78 resolved questions (59 Product Hunt, 19 Hacker News), 858 stated probability predictions, 175 full generation traces, June to August 2026. Models: Claude, GPT, and Gemini families, with and without web search. A Series A question family also ran but resolved only six times, which supports no conclusion, so it is excluded everywhere. Hacker News statistics cover 85 scored LLM predictions across 19 questions. Early numbers are stated as early numbers: samples are small, and definitive claims wait for the arena's scale.

Frequently Asked Questions

Can AI predict startup success?

The early signals are promising, and the definitive answer is being built. In Vela's two-month live arena, the most bullish calls from frontier models (GPT, Claude, Gemini) hit at 1.9 times the rate of chance, and the average of all models ranked outcomes better than any single model. These are early numbers on a small sample: a statistically definitive verdict needs roughly 800 resolved predictions, and the pilot collected 85. That is exactly what Vela's full arena is designed to collect, at scale, with every prediction scored against real outcomes.

What is an AI arena, like Arena AI, LMArena, or Design Arena?

An AI arena is a public competition where AI models take on the same task side by side and get ranked. Arena AI (arena.ai) and LMArena rank LLMs through millions of human preference votes on anonymous answers. Design Arena (designarena.ai) does the same for AI-generated designs. These arenas rank models by human preference: which answer people like. That works when taste is the ground truth. Venture needs a second kind of arena, one that checks whether a model's claims about the future come true.

How is Vela's VC Arena different from LMArena?

Preference arenas like Arena AI, LMArena, and Design Arena let humans vote on which output they like. Vela's VC Arena adds the dimension those arenas cannot: outcomes. Models state probabilities about live, unresolved events (a funding round, a launch, a repo's growth), the events resolve in the real world, and models are scored against what actually happened, including whether their probabilities were calibrated. Judgment tasks like market maps still use blind voting, but by verified investors and operators, and forecasts are graded by reality.

What is VCBench and what is the next version?

VCBench (vcbench.com) is the first AGI benchmark for venture capital, built by Vela Partners: 9,000 anonymized historical founder profiles with a public leaderboard. It tests models on frozen history. The next version, VC Arena, moves to the live future: models predict startup-world events before they resolve, so nothing can be memorized, and calibration is measured alongside accuracy. The two-month experiment described in this post was the design study for it.

What is VCRouter?

VCRouter (vcrouter.com) is the model router Vela is releasing alongside VC Arena. One API call sends any venture task, sourcing a thesis, forecasting a seed-to-Series-A conversion, drafting a market map, to whichever model currently leads that task in the arena. As the leaderboard moves, the routing moves with it. The design follows directly from Vela's arena data, where the blend of models outperformed every individual model.

Should you trust an LLM when it says it is 80% confident?

Not the raw number, not yet. In Vela's arena data, when models said an event was about 70% likely, it happened one time in seven, and models predicted success nearly three times more often than it occurred. The encouraging part: this is a calibration problem, not an intelligence problem, and calibration layers can be built and validated with enough resolved outcomes. That is one of the core things Vela's arena is built to measure continuously.

Does giving an LLM web search improve its predictions?

Not for forecasting, in Vela's paired tests. With and without web search, on the same questions, accuracy was identical to the percentage point for all three model families, while token cost rose by up to 32 times and responses became up to 4 times slower. For a question about the future, the answer is not on the web yet. Search buys background reading, not the answer.

Which LLM is best at forecasting: GPT, Claude, or Gemini?

On Vela's sample, the differences between frontier models were within noise, and the average of all models beat every individual model at ranking outcomes. That matches decades of forecasting research on human crowds: averaging beats champion-picking. It is the founding insight behind VCRouter, which blends and routes across models instead of betting on one.

How much does it cost to run LLM predictions at scale?

About a penny per prediction, or less. In Vela's arena, the 105 predictions with exact provider-reported bills cost $0.86 in total. The most expensive model cost three times the cheapest with no measurable accuracy difference, and the cheapest model was also the fastest. The economics of AI-assisted venture prediction already work; what the field needs now is resolved outcomes at scale.