Read every AI research paper. Listen to every podcast. Nod along at every dinner where someone explains why their portfolio company is "AI-native."

You still won't know what is possible versus what is not.

Not until you've tried to make these models do something hard yourself - and watched them fail, fluently, on your own data.

So we did. We built an AI system to score every deck that comes through 1752 - a multimodal pipeline that reads a deck the way an investor does, with a governance layer encoding how investors judge a company, calibrated against real decisions. Not to ship it. To find out what's real and what's BS - because you can't generate alpha from the cheap seats.

A few convictions came out the other side. Not predictions. Scars.

The Orchestration/Infrastructure Bet Is the Right One

Getting a multimodal model to do what you actually want - with meaningfully fewer hallucinations - is one hell of a lift.

We knew that abstractly. Now we know it personally.

The teams building those rails are solving the genuinely hard problem. Not the demo problem. The make-it-work-ten-thousand-times-in-a-row problem. The one that doesn't trend and doesn't fit in a keynote.

That's exactly why we want to back the rails. When the hype cycle exhales - and it always does - the companies holding up everyone else's product are the ones still standing.

Your Proprietary Data Is Your Partners' Judgment, Captured

You better have proprietary data to bend these models to your will. Your own, or sourced.

We used to say that as a thesis. Now I've lived it.

What's our proprietary data? It's how experienced investors actually evaluate a founder - the reasoning behind a yes, a no, a maybe. You can't buy that. It doesn't exist in a public dataset. We built it: we pulled signal from 25,000 real pitch decks and distilled it into a working set of more than three thousand principles the system scores against.

But raw extraction is just noise with ambition. The hard part - the part that actually took the work - was cleaning it. Cutting the duplicates. Killing the contradictions. Throwing out the principles that sounded sharp and said nothing. The signal of quality wasn't how much we pulled out. It was how much we were disciplined enough to cut.

That corpus is the moat. Not the model. Not the prompt. The captured judgment.

Everything else is rented.

AI-Native Beats Bolted-On - Every Time

Companies built around AI from day one outrun the ones retrofitting it. Not sometimes. Every time.

When AI is the foundation, the product and the data model are designed around what the tech can actually do. Bolt it on later and you get a feature stapled to a legacy architecture. Slower iteration. Shallower moats.

The proof is in our own build: the governance layer - the versioned rules for how we judge - isn't a setting buried in an admin panel. It's the spine. A competitor can call the exact same API we do. They can't call our constitution.

And here's what surprised me: you can see the difference within minutes of a demo. The AI-native product breathes. The bolted-on one strains - you feel the seams where the old system refused to bend. That gap isn't closing. It's widening.

The Calibration Gap Is the Work

Point the best frontier model at a stack of decks and ask it to score them like an investor would. It won't.

Out of the box, our system and our partners barely agreed. And when they diverged, the machine was almost always the optimist in the room - generous where a partner would've been brutal.

That gap is not a footnote. That gap is the product.

Most teams see early divergence, call it "directional," and ship. We measured it, stared at it, and closed it - slowly, through relentless calibration against real human reviews. The model gets you to the starting line. The loop is the race.

Clean data is the same fight wearing different clothes. One bad input and the system hallucinates - confidently. It doesn't fail loudly. It fails fluently. Nobody claps for a deduplicated corpus, which is exactly why most teams underinvest in the one thing separating a tool you trust from a liability that's wrong with a straight face.

Be Honest About What the Machine Actually Does

The AI scores the macro fundamentals well - high-level team analysis, problem, solution, stated market, financials. Human reviewers go deeper. They push past the stated views and ask the second-, third-, and fourth-level questions. Is that market size real, or backed into from a number the founder wanted to hit? Does the traction actually attribute to the product, or to a coincidence of timing? The judgment-heavy work still sits with people.

We didn't hide that line. We designed around it - and we keep moving it. As the calibration tightens, the AI takes on more. Today's limit is a floor, not a ceiling.

The temptation in this wave is to claim the machine does everything. The discipline is naming what it doesn't - yet. Stating the limit out loud is the only thing that makes the part it does do believable. Overclaim once and an insider stops trusting all of it.

Where the Value Actually Hides

Here's the lesson under all the other lessons.

The model labs are cleaning up everything deterministic, easy, and out in the open. Anything with one right answer - extract this field, summarize this doc, write this function - is getting absorbed into the frontier models themselves, for free, on their roadmap. If your startup lives there, you're a feature waiting to be deprecated.

So where does durable value go? It flees to three places the labs can't easily follow.

The non-deterministic. Evaluating a founder has no single right answer. Neither does scoring a deck, sizing a market, or calling a turnaround. It's judgment - contested, contextual, deeply human. The labs can't ship how a seasoned investor reads a founder - the instinct that separates a fundable team from a well-rehearsed one - because that lives in human heads, not in a benchmark.

The hard-to-reach data. A frontier model has read the entire public internet. It hasn't read what nobody publishes - the proprietary, the permissioned, the locked-up, the data you assemble one relationship and one painful integration at a time. The harder a dataset is to reach, the less likely a lab ever touches it. We didn't pull our 25,000 decks off a shelf. There is no shelf.

The genuinely hard problems. Not fiddly - hard. Years of domain pain, ugly edge cases, workflows no outsider understands. The labs optimize for the general case. The brutal, narrow, unglamorous problems are exactly the ones they skip - and exactly where a focused team builds something that lasts.

The through-line: the fuzziness, the inaccessibility, the sheer difficulty - everything the market treats as a bug - is the moat. The harder a thing is to make deterministic, to reach, or to solve, the longer it stays yours.

That's the counterintuitive part. In a world racing toward deterministic perfection, the defensible companies are the ones working exactly where it's hardest to follow.

Why We Put Ourselves Through This

The goal never changes. How do we create alpha for our LPs?

In this wave, you can't answer that from the cheap seats. The memos are downstream of someone else's conviction. The podcasts are downstream of someone else's agenda. Until you've tried to make these models do something genuinely hard yourself, you're trading on borrowed certainty.

Building it is how we earn the right to an opinion.

We didn't build this to become an AI company. We built it so that when we tell you something is real, we've already found out for ourselves - hands on the system, not ears on a recording of it.

Everyone else is reading about this wave.

We're building in it.