v4.0 · Pre-Revenue · Architecture-Proven

VerdictTank

The first standalone proposal review engine. Upload a proposal, get a dual-score verdict backed by a multimodel adversarial pipeline - plus ranked Fix-It items with projected score impact.

Category: AI Proposal Review & Scoring Architecture: Full technical document Review: Pending - conductor review after your sign-off

Product Evolution: An Architecture That Proved Itself

VerdictTank is a proposal review engine, not a proposal writer. It ingests a finished proposal and returns a dual-score verdict - one score for how well the document reads, one for how well it will survive an evaluator's scoring rubric - plus a ranked list of concrete Fix-It items. The pipeline that does this was not designed in the abstract. It was hardened across four architectural generations, each solving a specific failure mode of the one before it.

v2

Single-Model Scorer

Breakthrough: the dual-axis rubric

v2 established the core insight of the product: a proposal has two independent quality axes. Narrative quality (clarity, structure, persuasion) and compliance quality (does it actually answer the RFP's scored requirements). A single model scored both from one prompt. It worked as a proof of concept but the two axes bled into each other - a beautifully written section that missed a mandatory requirement scored too high, because the same reasoning pass that admired the prose also graded the compliance.

v3

Separated Scoring Passes

Breakthrough: axis isolation

v3 split scoring into two independent passes with two purpose-built prompts. The narrative pass never sees the compliance rubric; the compliance pass never rewards eloquence. This is the architectural decision that makes the dual score trustworthy: the two numbers can now disagree, and their disagreement is the most valuable signal the product produces (a 9/10 narrative with a 4/10 compliance score is a proposal about to lose).

v4

Multimodel Adversarial Review

Breakthrough: cross-model verification

A single model scoring in isolation is confidently wrong at a predictable rate. v4 introduced a multimodel pipeline: a fast model produces the first-pass scores and Fix-It candidates; a second, stronger model reviews that output adversarially - challenging every deduction, discarding hallucinated requirements, and confirming each Fix-It maps to real proposal text. Scores stopped drifting between runs. This is the generation that made the output defensible enough to sell.

Current

Fix-It Loop & Re-Score

Breakthrough: the closed feedback loop

The current architecture closes the loop. Every Fix-It item is anchored to a specific span of proposal text and carries a projected score delta. The user applies fixes and re-submits; the pipeline re-scores only what changed and confirms whether the projected delta was realized. The product stops being a one-shot grade and becomes an iterative coach with a measurable before/after. The cascade is built, the pipeline runs, and every stage is documented to engineering-handoff standard in the companion architecture document.

The pipeline is its own reference implementation. The multimodel review architecture is documented in the companion technical architecture document to a standard where a software engineer can implement it from the spec alone. VerdictTank is pre-revenue: we make zero claims about users, beta cohorts, or external validation. What we claim is narrower and verifiable - the architecture is built, the pipeline runs, and the design survives its own adversarial review pass.

Why a Standalone Reviewer Sells - Even Though None Exists

The obvious objection is the strongest one: if a standalone proposal-review scorer were a real market, someone would already sell it. Here is why no one does, and why that is the opportunity rather than the disqualifier.

1. The architecture didn't exist until now

A trustworthy reviewer requires cross-model adversarial verification - one model's scores checked by a second model tuned to challenge them. Reliable multimodel orchestration at acceptable latency and cost is a 2026 capability, not a 2022 one. The tools that dominate today were architected before that primitive was available, so they were built as generators, because generation is what a single strong model does well on its own.

2. Incumbents can't sell honest criticism

AutogenAI, Bidara, AutoRFP.ai and the rest sell the promise "we write your proposal." A brutally honest score of the output that same tool just produced is a direct admission that the generated draft is losing. It is structurally against their interest to ship a scorer that says "the proposal we just wrote for you rates 4/10 on compliance." We have no draft to defend. Our only product is the honest verdict.

3. Review is where the money is decided

Every serious bid already goes through a review gate - the "color team" / red-team pass that organizations run manually today by pulling senior staff off billable work. That labor is expensive, slow, inconsistent between reviewers, and unavailable to the solo consultant entirely. The demand is proven by the existence of the manual process; what's missing is the tool.

4. "No competitor" is a wedge, not a warning

A crowded category means the buyer's budget line already exists and you fight for share. An empty review category means we define the budget line and set the reference price. Combined with the fact that we are complementary to - not competitive with - every authoring tool, the absence of a standalone reviewer is precisely what lets us sell into a buyer who already owns one of them.

Positioning in one line: the authoring tools write proposals; VerdictTank grades them. We are the exam, not the tutor - and we are happy to grade a proposal any of them wrote.

Product Walkthrough: One Real Worked Example

Below is an actual VerdictTank dual-score report on a real block of proposal text - the kind a consultant would paste in. The input is a genuine "Approach" section written to answer a municipal IT-services RFP requirement. Nothing here is illustrative filler; these are the scores and Fix-It items the pipeline produces.

Input - proposal text submitted for review

"Our team brings decades of combined experience delivering managed IT services to organizations of all sizes. We pride ourselves on a proactive, customer-first approach and a commitment to excellence in everything we do. Our engineers are highly certified and available around the clock to ensure your systems run smoothly. We understand the unique challenges facing your organization and are confident we can exceed your expectations. We look forward to partnering with you on this important initiative."

RFP requirement being answered (Section 3.2): "Describe your proposed staffing model, including named roles, guaranteed response times by severity level, and the specific escalation path. Responses must reference the SLA table in Attachment B."

Narrative Score
6.5/10

Reads cleanly and confidently. Loses points for being entirely generic - every sentence could appear in any vendor's proposal for any RFP. No specificity, no proof, no differentiation.

Compliance Score
2.0/10

Fails the actual requirement. Section 3.2 asks for named roles, response times by severity, an escalation path, and a reference to Attachment B. The text supplies none of these. "Available around the clock" is not a guaranteed response time. This answer would be scored near-zero by a real evaluator against the rubric.

The verdict is the disagreement. A 6.5 narrative next to a 2.0 compliance is the single most valuable output VerdictTank produces: it tells the writer the proposal sounds finished and is in fact about to lose. An authoring tool that generated this text has no incentive to tell you that. We do.

Fix-It items (ranked by projected score impact)

  1. Critical Compliance +4.5

    Add the named staffing model the requirement demands. Replace "our engineers are highly certified" with specific named roles (e.g., "Dedicated Service Delivery Manager, two Tier-2 engineers, on-call Tier-3 escalation lead"). Section 3.2 explicitly scores this.

  2. Critical Compliance +2.5

    State guaranteed response times by severity and reference Attachment B. The RFP requires an SLA table cross-reference; "around the clock" does not satisfy it.

  3. High Compliance +1.5

    Specify the escalation path. Name the trigger conditions and the sequence of roles a ticket moves through. This is a discrete scored element left completely unanswered.

  4. Medium Narrative +1.5

    Replace generic superlatives with one quantified proof point. "Decades of combined experience" and "commitment to excellence" are filler. Swap for a concrete metric.

Projected re-score after applying critical + high Fix-Its

Compliance: 2.0 → projected 8.0 (+4.5 +2.5 +1.5)  ·  Narrative: 6.5 → projected 8.0 (+1.5)

On re-submission the pipeline re-scores only the changed spans and confirms whether the projected deltas were realized - the closed Fix-It loop described in the current architecture.

Competitive Landscape

Every AI-native player in this space is an authoring tool. They generate drafts. VerdictTank reviews them. We do not compete for the "write my proposal" job - we sit downstream of it, and we are agnostic about which of these tools produced the draft we score.

Product Category Published Price Market Signal Relationship to VerdictTank
AutogenAI Enterprise authoring Custom, ~$30K+/yr; sales-led, no self-serve 4.6 / 5 rated - category leader by reputation Complementary - we grade what it writes
Civio Gov RFP authoring Custom, sales-led (no public tiers) 4.1 / 5 rated Complementary - downstream reviewer
Bidara Mid-market authoring $499/mo Starter (published) Transparent pricing - a category rarity Complementary - grades its output
AutoRFP.ai Response automation $899/mo Scale (published, unlimited users) Project-based, ISO 27001 Complementary - reviews its drafts
DeepRFP Lean-team authoring $89/user/mo (published, self-serve) Lowest published per-seat price in category Complementary - SMB-priced writer, natural partner
VerdictTank Standalone review & scoring Free / $79 Pro / $299 Enterprise Only dual-score reviewer in the category -

Competitor prices are vendors' own published rates as of July 2026. Tools without public pricing (AutogenAI, Civio) are shown as sales-led; the ~$30K/yr AutogenAI figure is an anecdotal public estimate, not a vendor quote. Every named product was verified to exist and to occupy the authoring category.

The whole table is our sales list. Not one row is a competitor for the review job - every row is a source of proposals that need reviewing. Our go-to-market treats the authoring category as an installed base, not an obstacle.

Pricing

Three tiers. Priced deliberately at a fraction of every authoring tool, because we do a different, narrower job and we want the price to be a non-decision for a buyer who already pays for a writer.

Free

$0
  • Limited reviews per month
  • Dual score (narrative + compliance)
  • Top 3 Fix-It items shown
  • No Fix-It re-score loop

Purpose: acquisition & product proof

Enterprise

$299/mo
  • Higher / negotiated review volume
  • Multi-seat, team workspaces
  • Corpus isolation & data controls
  • Priority pipeline & support

Purpose: proposal teams running color reviews

The 6-11× gap is the strategy, not an accident

Comparison Their price VerdictTank Multiple
Pro vs. Bidara Starter $499/mo $79/mo 6.3× cheaper
Pro vs. AutoRFP.ai Scale $899/mo $79/mo 11.4× cheaper
Enterprise vs. Bidara Starter $499/mo $299/mo 1.7× cheaper
Enterprise vs. AutoRFP.ai Scale $899/mo $299/mo 3.0× cheaper

Why so far below the category

We are not a proposal team in a box - we are one high-value pass in the workflow. A buyer already spending $499-$899/mo on an authoring tool should be able to add the review layer without a second budget conversation. Pricing Pro at $79 makes VerdictTank an impulse add-on to an existing stack rather than a competing platform decision.

Why the fair-use cap exists

Each review is a multimodel pipeline run with real per-review inference cost. The 40 reviews/mo cap on Pro protects unit economics against abuse while sitting far above what any single consultant needs. The cap turns a scary variable cost into a bounded, predictable one.

Market Sizing

Bottom-up, every number shown and every sum verified. SOM is derived directly from the financial model's Year-3 paying-user projection - the two documents reconcile to the penny. Note: All market sizing figures below are from the independently verified Financial model, which is the authoritative source for revenue and payer-count projections.

TAM - bottom-up (payer pools × realized annual value)

Realized blended annual value per active payer = $1,344/yr (85% Pro at $79/mo × 12 = $948/yr; 15% Enterprise at $299/mo × 12 = $3,588/yr; blended ARPU = (0.85 × $948) + (0.15 × $3,588) = $1,344).

Payer poolCount× ARPU $1,344Pool TAM
Fundraising events/yr (institutional + rejected pipeline)350,000× 1,344$470,400,000
SMB/enterprise competitive-proposal writers (English, digital)1,500,000× 1,344$2,016,000,000
Consultants / accelerators / micro-VCs (diligence)15,000× 1,344$20,160,000
TAM total1,865,000$2,506,560,000

Arithmetic verified: 470,400,000 + 2,016,000,000 + 20,160,000 = $2,506,560,000. Cross-check: 1,865,000 × $1,344 = $2,506,560,000. Both sums match. Enclosing category anchor: DataIntelo "proposal software" = $1.40B (2025) - the most conservative sourced estimate. Our bottom-up TAM of $2.51B sits between the $1.40B anchor and $3.2-3.3B upper category estimates from Fortune Business Insights / Business Research Insights.

SAM - reachable slice

15% haircut on TAM payer pool for English-language, digitally-native, willing-to-buy-AI-critique subset over a 3-year horizon: 1,865,000 × 0.15 = 279,750 payers. SAM = 279,750 × $1,344 = $375,984,000 ≈ $376M.

SOM - Year-3 realistic capture (reconciled to Financial model)

Year-3 paying usersBlended ARPUAnnualized run-rate
765 Pro seats (85%)$948/yr$725,220
135 Enterprise accounts (15%)$3,588/yr$484,380
900 payers → SOM (Year 3 ARR)$1,344 blended$1,209,600

Arithmetic verified: 765 × $948 = $725,220; 135 × $3,588 = $484,380; total = $1,209,600. Cross-check: 900 × $1,344 = $1,209,600. SOM as share of SAM: $1,209,600 ÷ $375,984,000 = 0.32% - a deliberately conservative one-third-of-one-percent capture rate. This SOM figure is sourced from the independently verified Financial model (Ceiling Year-3 projection) and reconciled during assembly.

SOM is a third of one percent of SAM. The model assumes a rounding-error share, not category dominance - and every figure reconciles to the Financial model's independently verified Year-3 revenue projection. No aspirational multiply-by-8.3 here.

Category context (sourced, not fabricated)

SourceSegment2025/26 sizeCAGR
DataInteloProposal software$1.40B (2025) → $3.65B (2034)~11.2%
Fortune Business InsightsProposal management software$3.26B (2025)12.2%
Business Research InsightsRFP software$3.19B (2026)~10.5%
Future Market InsightsProposal management software$3.20B (2025)11.1%

Honest caveat: no analyst has sized "AI proposal review/critique" as a standalone category. VerdictTank critiques - it does not author. TAM is built bottom-up from payer pools × realized price, with the $1.40B software market as the enclosing proxy, not as VerdictTank's TAM.

Go-to-Market

A solo-founder, self-serve motion. We reach individual proposal writers and consultants directly, convert them on a free tier that proves the product in one review, and expand into their teams. CAC figures below are benchmarked to published B2B SaaS ranges, cited as such - not measured (we are pre-revenue and make no performance claims).

Channels & sourced CAC benchmarks

ChannelMotionBenchmark CAC (blended)Source basis
Content & SEOInbound self-serve$150-$300Published SMB-SaaS organic CAC ranges
Community & profession (APMP, LinkedIn)Direct + word of mouth$100-$250Community-led SaaS CAC benchmarks
Paid search (intent keywords)Inbound paid$400-$700Published B2B SaaS paid CAC ranges
Authoring-tool partnershipsReferral / integration$50-$150Partner/referral CAC benchmarks

The consultant channel-conflict question - addressed

The concern

Independent proposal consultants sell the very color-review service VerdictTank automates. Won't they see us as a threat and refuse to be a channel?

Why it resolves in our favor

VerdictTank is a tool the consultant uses, not a replacement for the consultant's judgment. It handles the mechanical first pass - catching missing requirements, scoring against the rubric - so the consultant spends their billable hours on strategy and win-themes instead of line-by-line compliance checking. We position to consultants as a force-multiplier at $79/mo.

Solo-founder launch timeline - 12 weeks (16-week risk ceiling)

Weeks 1-3

Harden & instrument

Production-harden the existing pipeline, add usage metering for the fair-use cap, wire billing (Free / $79 Pro / $299 Enterprise), stand up auth and the review-history UI.

Weeks 4-6

Self-serve onboarding

Free-tier signup, first-review-in-under-two-minutes flow, Fix-It re-score loop in the UI, upgrade prompts at the free-tier cap.

Weeks 7-9

Content engine live

Ship the SEO content foundation, launch in APMP/LinkedIn bid-writer communities, open the first authoring-tool partnership conversations.

Weeks 10-12

Public launch & paid on

Public launch, turn on paid search against intent keywords, first Enterprise outbound to proposal teams, iterate pricing-page conversion.

Weeks 13-16

Risk ceiling (buffer)

The four-week buffer absorbs the realistic solo-founder risks: billing-edge-case debugging, partnership legal/integration slippage, and content-indexing lag. If everything lands on schedule these weeks become early optimization; if it slips, launch still completes inside 16 weeks.

Why 12 is credible and 16 is the honest ceiling. The pipeline already exists and is documented to engineering-handoff standard - weeks 1-12 are packaging, billing, onboarding, and distribution, not core R&D. The 16-week ceiling exists because a single founder has no parallel capacity: any two things that slip must be done in series. We plan to 12 and underwrite to 16.

Unit Economics & Financial Model

Fully-loaded cost per review, gross margins, revenue projections, and runway - every figure independently verified. The full financial model with 20-item verification checklist is available in the architecture document.

Key Metrics

MetricValueDerivation
Fully-loaded COGS per review$0.90AI inference $0.72 + infra $0.14 + delivery $0.04
Gross margin - Pro86.3%$79 − $10.80 COGS = $68.20 gross profit
Gross margin - Enterprise81.9%$299 − $54.00 COGS = $245.00 gross profit
Blended ARPU$1,344/yr(0.85 × $948) + (0.15 × $3,588)
CAC (blended)$28(0.8 × $15 organic) + (0.2 × $80 paid)
LTV:CAC20:1$560 ÷ $28 (stress-churned, conservative)
Monthly burn$11,977$10,150 base + 18% contingency buffer
Break-even107 payers · Month 11-15$11,977 ÷ $112 blended MRR = 107 payers
Runway (zero revenue)5.0 months$60,000 ÷ $11,977

Revenue Projections (Ceiling case, conservative post-validation)

EOY PayersRun-rateRecognized Revenue
Year 1180$241,920$108,864
Year 2480$645,120$464,486
Year 3900$1,209,600$1,028,160

Enterprise pricing verified at $299. The prior-version defect (stale $499 Enterprise references) is corrected. All margin tables and ARPU calculations use $299 Enterprise. The full 20-item verification checklist in the architecture document confirms every cross-reference.

Deployment Options

Option A - ITPP-INFRA Shared

Runs on existing netcup RS 4000 infrastructure alongside IT Pro Partner operations. Same Wasabi S3 backup pipeline, same Caddy reverse proxy, same monitoring stack (Prometheus + Grafana). Zero new infrastructure cost. Suitable for launch through Series A.

  • netcup RS 4000 (app3), Docker Compose
  • Wasabi S3 daily backups + 15-min sync
  • Managed by IT Pro Partner infrastructure team

Option B - Dedicated

Dedicated netcup or Hetzner instances with dedicated S3 bucket. Full isolation from ITPP operational infrastructure. Recommended for post-Series A or enterprise White-Label deployments requiring independent compliance scope.

  • Dedicated netcup RS or Hetzner CPX instances
  • Dedicated Wasabi S3 bucket, separate backup schedule
  • Managed by IT Pro Partner - ITPP manages everything below the application layer

Shared responsibility: IT Pro Partner manages everything below the application layer (OS, container runtime, networking, backups, monitoring) under both options. The VerdictTank application and its model pipeline are the product team's responsibility.