Today’s newsletter is sponsored by Xplor Pay.
Artificial intelligence is changing how software is built, customer expectations are evolving, and investors are rethinking what creates long-term enterprise value.
Join Xplor Pay and Luke Sophinos, Founder of Vertical SaaS Group, for an executive discussion exploring how AI, workflow ownership, platform strategy, and financial capabilities are reshaping the future of vertical SaaS. Register today and reserve your spot.
Vertical-Specific GPT’s
A lot of vertical AI companies are not losing to incumbents.
They are losing to ChatGPT.
The founders who understand this are quietly building something different. They are not shipping features. They are shipping answers. Better, deeper, more specific answers than any horizontal model can produce, for the exact questions their customer asks every single day.
If your vertical AI product cannot out-answer GPT on the top 1000 questions your customer actually asks, you do not have a moat. You have a wrapper with a login screen.
This week is about how to actually build that moat, why the Vertical Software Summit is the only room where operators are talking about it honestly, and which vertical GPTs are already proving the model works at scale.
Prompting Your Way to an Epic Vertical AI Use Case
Most vertical AI products die at the prompt layer. The team ships a chat interface, wraps a horizontal model, adds a domain-specific system message, and calls it a product. Six months later the same customer asks the same question in ChatGPT and gets a comparable answer. The wrapper is exposed. Retention collapses. The team blames distribution.
Distribution is not the problem. The answer is the problem.
Here is what the operators actually winning in vertical AI are doing. This is the full playbook.
Step 1: Map the top 1000 questions in order of importance.
Not 10. Not 100. One thousand.
The number is not arbitrary. Below 100 you have a demo. Below 500 you have a feature. At 1000 you have a product. Above 1000 you start covering the long tail where retention actually lives.
How to build the list:
Pull two years of support tickets. Every ticket is a question in disguise.
Transcribe every sales call, onboarding call, and QBR from the last 12 months. Extract the questions the customer asked, not the ones your team answered.
Sit inside the customer for a full week. Not a demo day. A real operating week. Watch what they Google. Watch what they ask their coworker. Watch what they paste into ChatGPT when they think nobody is looking.
Scrape the industry subreddits, the private Slack groups, the closed Facebook groups, and the trade forums. That is where the real questions live.
Interview 25 customers with one prompt: what did you ask an expert this week that you wish software could answer.
Then rank them. Two axes. Frequency and economic weight. A question asked 40 times a week by a billing manager beats a question asked once a quarter by the CFO. Order the list. That order is your build sequence.
This list is not a document. It is your product roadmap. Everything downstream of this list is execution.
Step 2: Build a massively detailed prompt for each question.
Twenty pages is not a joke. It is the floor.
The best vertical prompts in production today are structured like operating manuals written by a senior domain expert who is also a systems thinker. Every one of them contains, at minimum:
Domain framework. The mental model an expert uses to approach this question. Not the answer. The way of thinking.
Vocabulary layer. The exact terms of art, the acronyms, the regional variations, and the terms that mean different things in different sub-verticals.
Regulatory context. What the customer is legally allowed to do with this answer. What disclaimers must appear. What jurisdictions apply.
Workflow context. Where in the customer’s day this question gets asked. What tool they are in when they ask it. What the answer needs to plug into next.
Output format. Text, table, filled form, diagram, deck, doc, cited citation list. Different questions want different artifacts.
Edge cases. The top 20 ways this question can be asked incorrectly, incompletely, or with hidden assumptions.
Graceful degradation. What the answer looks like when 30 percent of the required data is missing. What it looks like at 60 percent missing. What triggers a human handoff.
Source hierarchy. Which sources are authoritative, which are supporting, which are disqualified.
Anti-hallucination rules. The specific claim types this prompt is not allowed to make without a citation.
Tone and register. How a senior person in this vertical actually talks.
If your prompt fits on one screen, you have not done the work. You have written a system message.
Step 3: Test every model against every prompt.
GPT, Claude, Gemini, open source frontier models, image models, deck generators, doc generators, spreadsheet generators, form fillers. Do not have a favorite model. Have a favorite outcome.
For every question in the 1000, run the prompt through every viable model. Score the output on:
Factual accuracy against a domain-expert answer key.
Format fidelity to what the customer actually needs.
Latency at the point of use.
Cost per answer at scale.
Failure mode when the input is broken.
Some questions want Claude for reasoning. Some want GPT for structured output. Some want Gemini for long context. Some want an image model to produce an annotated diagram. Some want a deck generator to produce a client-ready output in 30 seconds. Match the model to the answer, not to the vendor relationship.
The customer does not care which model ran. The customer cares that the answer was right.
Step 4: Benchmark every answer against horizontal AI.
For each of the 1000 questions, paste the raw question into ChatGPT, Claude, and Gemini as a naive user would. Capture that answer. That is your baseline.
Your product must produce an answer that is 10x better than the baseline. Not 20 percent better. Not 2x better. 10x.
Ten times better means one or more of the following:
Cited where the baseline is not.
Structured for the workflow where the baseline is prose.
Compliant where the baseline is generic.
Current where the baseline is stale.
Personalized to the customer’s data where the baseline is universal.
Delivered inside the tool where the baseline requires a copy paste.
If you cannot articulate the 10x on a per-question basis, that question is not yet a moat. It is a coin flip. Coin flips lose to whoever has the better distribution, and horizontal AI has better distribution than you do.
Step 5: Build the master spreadsheet.
This is the artifact that separates operators who talk about vertical AI from operators who ship it.
Columns:
A: Question ID
B: The question in the customer’s actual words
C: Ranked importance (1 to 1000)
D: Frequency per customer per month
E: Economic weight (dollars touched by this question)
F: The full prompt (link to the 20 plus page doc)
G: All similar questions that route to this same prompt (typically 5 to 30 variants)
H: The winning model for this question
I: Output format
J: Latency target
K: Cost per answer
L: Horizontal AI baseline answer (link)
M: Your answer (link)
N: The specific 10x delta (one sentence)
O: Score against baseline (1 to 10)
P: Last tested date
Q: Owner
R: Retest cadence
That spreadsheet is your real product spec. The UI is downstream of it. The pricing is downstream of it. The GTM is downstream of it. The moat is inside it.
Step 6: Retest on a cadence.
The horizontal models get better every quarter. Your 10x lead compresses if you do not defend it. Set a mandatory retest cadence. The top 100 questions retest monthly. The next 400 retest quarterly. The bottom 500 retest twice a year. If any question drops below 10x against the current frontier, it moves back into the active build queue.
This is not a project. It is an operating discipline.
The uncomfortable truth: being in the workflow is not a moat anymore. Answering domain questions better than any horizontal model, on a per-question basis, benchmarked and defended, is the moat. If a horizontal model can match your answer, your customer will eventually notice, and your product will collapse into a login screen on top of someone else’s intelligence.
The work is boring. The 1000 questions are boring. The 20 page prompts are boring. The spreadsheet is boring. That is exactly why it compounds. Nobody else is willing to do it.
The Vertical Software Summit 2026
There is a reason we built this thing.
For a decade the vertical software conversation has been fragmented. Scattered X threads. A handful of newsletters. One or two tracks buried inside larger horizontal conferences. A lot of private group chats and off-the-record dinners. The category that quietly produces most of the durable software value in the world had no home.
The Vertical Software Summit is that home.
What makes the format work is what it refuses to be.
It is not a payments conference masquerading as a vertical software confernece.
It is not an investor conference not allowing other investors in.
It is the only conference by vertical software and ai founders for vertical software and ai founders.
The Summit is founders, operators, investors, and all the companies that support vertical businesses talking to each other, in one room, without the noise.
The lineup tells the story:
Founders who built systems of record in industries most tier one VCs would not touch. Pool cleaning software. Woodworking software. Pizza shop software. Laundromat software. The unglamorous categories that produce 40 percent EBITDA and 110 percent net revenue retention.
Acquirers & Investors who understand that a boring vertical with those numbers is worth more than a horizontal darling burning cash at scale.
Operators who have survived the SaaS Crash, the SVB bank crisis, the SaaSpocalaypse and are now walking through the software to AI transition in real time. This is the room where founders learn what the second transition actually costs, whats working, from the people currently paying for it.
AI-Native companies that are showing us how fast growth can look when you launch a net-new vertical AI product.
The programming is centered around:
War stories and hard lessons learned from the folks that have monopolized their verticals.
Playbooks we can all take into our own businesses that are already proving to work in adjacent industries.
CxO’s who can go deep in particular areas — AI-powered sales enablement, GTM, Customer Success, etc.
The bet underneath the event is simple. Vertical software is not a niche of the tech industry. It is the majority of durable software value creation. The market has priced this in at the incumbent level. It has not priced it in at the emerging company level. That gap is where the next decade of returns lives.
The people already in the room know this. The Summit is where they compare notes.
If you build, invest in, or acquire vertical software companies, this is the room you should be in. I hope to see you there! Grab your ticket, come enjoy Miami in November, and do it before tickets sell out!
The Vertical GPTs That Are Actually Winning
The fastest growing AI companies in the world right now are not horizontal chat products. They are vertical-specific GPTs that decided, early, to out-answer the horizontal models on a narrow, high-value domain. Every single one of them, whether they call it that or not, has been running some version of the 1000 question playbook.
Harvey did it for elite law firms first. The product is not a chat interface. The product is an internal framework that reflects how BigLaw associates actually draft, review, cite, and defend legal work. Horizontal models produce plausible legal text. Harvey produces text a senior partner will sign. That difference is worth every dollar the top firms pay for it, and it is why Harvey now sits inside a large share of the AmLaw 100. The moat is not the model. The moat is the encoded understanding of how legal work actually gets produced inside a firm, and the citation and review infrastructure that makes the output defensible.
MagicSchool did it for K-12 teachers. The horizontal models can technically write a lesson plan. What they cannot do is write a lesson plan that maps to a specific state standard, differentiates for three reading levels in the same classroom, aligns with the district’s approved curriculum, and outputs in the exact format the teacher’s LMS expects. MagicSchool encoded that specificity across dozens of daily teacher workflows and became one of the fastest adopted education products of the AI era, now used by millions of educators across tens of thousands of schools. Teachers do not stay because the model is better. They stay because every answer lands inside their actual workflow.
OpenEvidence did it for practicing physicians. Founder Daniel Nadler had already sold Kensho to S&P for around 700M. He did not need to build this. He built it because he saw that the highest frequency, highest stakes questions in American medicine were being answered either by memory, by outdated UpToDate articles, or by physicians pasting patient scenarios into ChatGPT under the desk. OpenEvidence mapped those questions. Cited every answer to peer-reviewed medical literature. Made it free at the point of use. Monetized, quietly, through embedded pharma advertising to a captive audience of prescribers. The result is one of the fastest ARR ramps in software history, roughly 150M in run rate, reached faster than almost any software company on record. A meaningful percentage of practicing US physicians now use it daily. The horizontal models cannot compete on this domain because they were never given the citation infrastructure, the regulatory awareness, or the workflow context that OpenEvidence encodes by default.
ChipAgents did it for semiconductor design. This is the vertical where horizontal AI fails most visibly. The vocabulary alone requires years of encoded context. RTL, verification, timing closure, DFT, tapeout. A general purpose model produces confident nonsense the moment you push past surface level questions. ChipAgents encoded the domain, integrated with the EDA toolchain, and gave semiconductor engineers an answer layer that actually understands what they are asking. The company raised 21M to keep expanding that answer surface, and its wedge tells you exactly where vertical AI has structural advantage: the more specialized the vocabulary and the more expensive the mistake, the wider the moat.
TrunkTools did it for commercial construction. Construction is the perfect vertical AI target and the perfect vertical AI graveyard. Every project generates thousands of pages of specifications, drawings, RFIs, submittals, and change orders. Every question a superintendent asks in the field has a definitive answer buried somewhere in that document set, and no horizontal model can find it because horizontal models were not trained on a specific project’s document universe, do not understand the specification hierarchy, and cannot cite the exact page a general contractor needs to defend a decision. TrunkTools built the answer layer on top of that document set. A super in the field asks a question. TrunkTools returns the exact spec section, the exact drawing detail, the exact RFI response, cited to the page. The 1000 questions in construction are things like: what is the required fire rating for this partition, what submittal governs this fixture, what is the latest revision of this drawing, which RFI answered this clash. Horizontal AI cannot touch any of that. TrunkTools can, and the general contractors and subcontractors adopting it are compounding a data advantage horizontal models will never replicate.
The pattern under every one of them is identical. They did not win because they had access to a better base model. They won because they took the boring, unglamorous work of mapping their customer’s real questions, building the answer infrastructure around them, and defending that infrastructure with citations, workflow integration, and continuous benchmarking against the horizontal frontier.
The horizontal models are impressive. The vertical products are indispensable.
That is the difference between a demo and a business, and it is the difference between an AI company that will still exist in 2030 and one that already quietly does not.
See you next Sunday.
Luke
Do me a solid and forward to a friend :-)










