← Back to Proposals
Confidential - Advisory Review

RFP Tank

AI-powered government contract proposal platform - find, analyze, and win federal contracts. Scores a proposal against the RFP for compliance, then tells you whether it will win.

📅 August 2026🏢 IT Pro Partner - Product Division🌐 rfptank.com

Marketing Site

rfptank.com
Not created

Dashboard/App

app.rfptank.com
Not created

API

api.rfptank.com
Not created

Domain

rfptank.com
Purchased

1. Business Proposal

Full product proposal - business case, differentiation, pricing, go-to-market, financials, and the ask.

Prepared by: IT Pro Partner | Germaine Brown

Date: August 2026

Version: 1.0


"Throw your RFP in the Tank. Out comes a winning response."

  1. Executive Summary
  2. Elevator Pitch
  3. Problem Statement
  4. Market Analysis
  5. Product Overview
  6. Revenue Model
  7. Competitive Advantages (Moat)
  8. Go-to-Market Strategy
  9. Risk Analysis
  10. Financial Projections
  11. The Ask

1. Executive Summary

The United States federal government spent $793 billion on contracts in FY2025. Every single dollar of that required someone to write a proposal. That proposal funnel is the choke point for an entire industry - and the tools serving it are broken.

The enterprise players hide pricing behind a sales call and target mid-to-large primes. The cheap tools are search wrappers with a chatbot bolted on. The middle - the 200,000+ active small businesses, independent consultants, and emerging contractors who collectively chase $182 billion in set-aside contracts every year - gets nothing purpose-built. They get spreadsheets, prayer, and $500/hour consultants they cannot afford.

RFP Tank fills that gap with a sledgehammer.

RFP Tank is a SaaS platform purpose-built for government contractors who need to find, analyze, and respond to RFPs faster and more accurately than any tool on the market. It is not a search tool with AI added. It is not an enterprise platform with a self-service tier bolted on. It is a complete proposal pipeline - from proactive opportunity matching to guided response flow to compliance checking to past performance mining - built from the ground up for the contractor who needs to WIN, not just submit.

The differentiation is twofold. First, depth on traceability: a compliance matrix that maps every mandatory requirement to the exact paragraph that answers it, with a confidence score and quoted evidence. Second, an evaluation layer no competitor has: an evaluator simulation that scores the proposal the way the government will, using the RFP's own criteria and weights, so the contractor knows whether it will win, not just whether it complies.

Market Reality:

Revenue Potential (Year 1 Conservative):

The competitive comparison is not close. GovEagle targets mid-to-large primes - no public pricing, no self-service, enterprise sales cycle only. CLEATUS is a SAM.gov wrapper. GovDash competes on price, not quality. Nobody has built the full guided proposal workflow for self-service contractors. Nobody. Until now.

RFP Tank is an IT Pro Partner product. The infrastructure is already running. The SAM.gov API is free. The USASpending.gov data is public. The value is in the processing, the workflow, and the intelligence layer - and that is exactly what we are building.

The Ask: Domain (rfptank.com - confirmed available), 90 days of focused development, and the confidence to price like the best product in the market - because that is what it will be.


2. Elevator Pitch

GovEagle's pricing is unpublished and getting in the door requires an enterprise sales cycle. The other 99% of government contractors - the small businesses, independent consultants, and emerging prime candidates chasing $182 billion in set-aside contracts - are stuck writing proposals in Microsoft Word with a SAM.gov tab open and a spreadsheet for their compliance matrix.

RFP Tank is GovEagle for the other 99%.

Upload any RFP. In 60 seconds, get a complete compliance matrix, requirement breakdown, and Section L/M alignment. Then work through a guided response wizard that checks every answer against every requirement in real time. Build your past performance library once, mine it forever. Track every bid. Learn what wins. Before you submit, run an evaluator simulation that scores the proposal the way the government will, so you know whether it will win, not just whether it complies.

This is not a feature gap between us and the competition. This is market failure at $793 billion scale - and RFP Tank is the correction.


3. Problem Statement

3.1 The Broken Status Quo

Government contracting is a $793 billion market with a proposal problem. Every awarded contract requires a winning response to a solicitation. The average federal RFP runs 50-200 pages. It contains dozens of mandatory requirements, multiple evaluation criteria, a compliance matrix that must be manually built, and formatting rules that disqualify non-conforming submissions. A single RFP response takes 40-200 hours of work. A mid-sized contractor might pursue 20-30 per year.

Here is how contractors currently handle this:

Current Method Who Uses It Time Cost Dollar Cost Pain Level
Microsoft Word + SAM.gov tab Most small contractors 60-200 hrs/RFP $0 tool cost, $15K-$50K labor Extreme
Proposal consultants Mid-sized firms 40-100 hrs/RFP $500-$1,500/hr High + expensive
GovEagle / enterprise tools Mid-to-large primes (50+ employees) 20-60 hrs/RFP Unpublished (enterprise) Low tool friction, high cost
CLEATUS / search wrappers Individual contractors 40-100 hrs/RFP $39-$78/mo Moderate discovery, no response help
Spreadsheet compliance matrix Everyone 5-20 hrs per RFP $0 Tedious, error-prone, manual

The result: small contractors are fighting $793 billion in opportunity with tools designed for a different era. They lose not because their work is inferior - they lose because their proposals are non-compliant, incomplete, or poorly structured. The requirement was there. They missed it. Or they addressed it halfway. Or they forgot to map it to an evaluation criterion. The government evaluator marks them down. The work goes to someone with a better proposal team.

3.2 The Target User

RFP Tank serves three distinct users, all underserved:

Persona 1: The Solo Contractor

Persona 2: The Small GovCon Firm

Persona 3: The Emerging Prime

3.3 The Gap That RFP Tank Fills

The market is not missing a search tool. SAM.gov is free. The market is missing a proposal execution engine - a tool that takes you from "I found an opportunity" to "I submitted a winning, compliant response" with AI assistance at every step. No existing tool does this for self-service users at any price point. That is the gap. It is not subtle. It is not a feature gap. It is a complete market absence at the most critical point in the contractor workflow.


4. Market Analysis

4.1 Total Addressable Market (TAM)

The federal government contract market is one of the largest and most consistent spending pools in the world. It does not contract in recessions. It does not go to zero. It carries a statutory 23% small business contracting goal, roughly $182 billion of the $793 billion total.

Market Segment Annual Spend Relevance to RFP Tank
Federal contract spending (FY2025) $793 billion Every dollar flows through a proposal
DoD contracts alone $398.2 billion (50.2%) Largest single agency pool
Small business set-asides $182 billion (23% statutory goal) Primary target user base
SLED (State, Local, Education) ~$1.6 trillion 2x the federal market, more accessible
Cooperative purchasing market $65+ billion Accessible via GS schedules and co-ops
Contracts expiring in 18 months $142 billion (85,247 contracts) Recompete pipeline - highest-value targets

TAM: Roughly 200,000 active small business government contractors (per the SAM analysis in 4.2). At a $200/month average for a dedicated proposal tool, the addressable market is $480 million ARR.

Capture target: Even capturing 2% of the $480M TAM is $9.6M ARR - a venture-scale business from a small customer base. The near-term target (619 accounts, 0.31% penetration in 4.3) is deliberately conservative against this ceiling.

The key insight: A 1% win-rate improvement on the small business set-aside portion alone ($182 billion) is worth $1.8 billion in additional contract awards to small businesses. If RFP Tank helps a contractor win one additional $500K contract per year, it has delivered 50x the value of a $599/month subscription. That is the unit economics story that makes this product a no-brainer purchase.

4.2 Serviceable Addressable Market (SAM)

User Type Estimated Count RFP Tank Relevant Serviceable
Active small business contractors (solo, firms, and emerging primes) 200,000 100% relevant 200,000
Total SAM ~200,000

The 200,000 is an umbrella spanning independent consultants, 2-20 person firms, and emerging primes of up to 100 employees - all of which qualify as small businesses under SBA size standards, counted once to avoid double-counting. The 500,000+ registered SAM.gov entities is the top-of-funnel universe; roughly 200,000 are active pursuers.

4.3 Serviceable Obtainable Market (SOM)

Year Target Subscribers ARR Penetration of SAM
Year 1 619 $1.28M 0.31%
Year 2 2,000 $5.1M 1.00%
Year 3 6,000 $16.2M 3.00%

Year 1 penetration of 0.31% requires finding 619 customers in a market of 200,000. That is not a growth problem. That is a targeting problem - and GovCon is one of the most targetable communities in B2B software (LinkedIn by NAICS code, GovCon forums, APEX Accelerators, 8(a) cohort lists).

4.4 Competitive Landscape

The competitive comparison is not close. It is embarrassing for the incumbents:

Feature RFP Tank GovEagle GovDash CLEATUS GovSignals Vultron Rohirrim Proc. Sciences
Self-service signup YES NO NO YES NO NO NO NO
Transparent pricing YES NO NO YES NO NO NO NO
Profile-driven RFP matching YES NO NO Partial NO NO NO NO
RFP shredder (60-sec compliance matrix) YES YES YES NO YES YES YES YES
Guided step-by-step response flow YES NO NO NO NO NO NO NO
Real-time compliance check per answer YES NO NO NO NO NO NO NO
Past performance vault + AI mining YES Enterprise only NO NO NO NO NO NO
Win/loss analytics YES NO NO NO NO NO NO NO
Solo/individual tier YES NO NO YES NO NO NO NO
Team collaboration YES YES YES NO YES YES YES YES
Serves solo + SMB + mid-market YES NO Partial Partial NO NO NO NO
Walk-up entry price $79/mo Enterprise Custom $39/mo Enterprise Custom Custom Custom
Team-level price $599/mo Unpublished Custom $78/mo Custom Custom Custom Custom

Nobody does all of this. Nobody. Until now.

RFP Tank column reflects planned MVP capability at launch. Build status by feature is tracked in Section 5.

4.5 Why Competitors Cannot Simply Copy This


5. Product Overview

5.1 Architecture

+------------------------------------------------------------------+
|                     RFP TANK PLATFORM                            |
+------------------------------------------------------------------+
|                                                                  |
|   DATA INGESTION LAYER                                           |
|   +-----------+  +--------------+  +------------------+         |
|   | SAM.gov   |  | USASpending  |  | State Procurement|         |
|   | API (free)|  | API (free)   |  | Portals (scraped)|         |
|   +-----------+  +--------------+  +------------------+         |
|         |               |                   |                   |
|         +---------------+-------------------+                   |
|                         |                                        |
|   AI PROCESSING PIPELINE                                         |
|   +-----------------------------------------------------+        |
|   | Opportunity Scorer | RFP Shredder | Content Miner  |        |
|   | (profile match)    | (60-second)  | (past perf)    |        |
|   +-----------------------------------------------------+        |
|                         |                                        |
|   USER LAYER                                                     |
|   +------------+  +------------------+  +------------------+    |
|   | Smart      |  | Guided Response  |  | Past Performance |    |
|   | Discovery  |  | Flow + Compliance|  | Vault + Mining   |    |
|   | Dashboard  |  | Copilot          |  |                  |    |
|   +------------+  +------------------+  +------------------+    |
|                         |                                        |
|   OUTPUT LAYER                                                   |
|   +----------+  +----------------+  +------------------------+  |
|   | Word/PDF |  | Compliance     |  | Win/Loss               |  |
|   | Export   |  | Matrix (Excel) |  | Analytics Dashboard    |  |
|   +----------+  +----------------+  +------------------------+  |
+------------------------------------------------------------------+

Technology stack: Node.js/Python backend, React frontend, PostgreSQL for user data, vector database for past performance content retrieval, LLM API for compliance checking and content generation.

5.1.1 Deployment Options - ITPP-INFRA Shared vs Dedicated

This section addresses the single most common enterprise buyer question: "where does this run?" RFP Tank is not an abstract cloud service - it runs on bare-metal infrastructure owned and operated by IT Pro Partner, with two deployment models to match the client's risk and isolation requirements.

Option A: ITPP-INFRA Shared Deployment (Default)

The solution runs on IT Pro Partner's existing production infrastructure (ITPP-INFRA), the same hardware and backup pipeline that powers ITPP's own operations.

Layer Hardware Details
Edge netcup RS 4000 (app3) or Hetzner CPX (app1-bu) Caddy TLS termination, domain routing, rate limiting
App netcup RS 4000 (app2, app3) or Hetzner CPX21 Containerized services (Docker), autoscaling within the pool
Data netcup RS 4000 (app2) PostgreSQL 16 for user data + vector DB for past performance embeddings
Object Storage Wasabi S3 (shared bucket) Proposal archives, uploaded RFPs, generated response documents
Backup ITPP backup pipeline Daily S3 sync, versioning ON, 90-day retention; database WAL archiving via hermes-live-sync

Best for: Solo tier, Consultant tier, early Team tier pilots. Cost-efficient - leverages existing capacity with no dedicated hardware overhead. All infrastructure is documented and auditable under ITPP's existing operational controls.

Option B: Dedicated Deployment

The solution runs on US-owned, US-hosted dedicated infrastructure - AWS GovCloud, Azure Government, or a US-owned colo - provisioned specifically for the client, with a dedicated object-storage bucket and isolated backup pipeline. This tier exists specifically for workloads that cannot run on shared infrastructure: CUI, ITAR-controlled data, or covered defense information under DFARS 252.204-7012.

Layer Hardware Details
Edge Dedicated AWS GovCloud or Azure Government Isolated reverse proxy, client-specific TLS
App Dedicated US-owned compute (GovCloud / Azure Gov) No shared compute with other RFP Tank tenants
Data Dedicated PostgreSQL + vector DB on the same instance or separate, per client requirements Client data never co-mingles with other tenants
Object Storage Dedicated US-region object storage (Wasabi US or GovCloud S3) No cross-tenant object storage; bucket owned by client or ITPP per contract
Backup Dedicated backup pipeline Independent S3 bucket with versioning ON; dedicated cron schedule; restore testing included quarterly

Best for: Enterprise clients handling CUI or ITAR-controlled solicitations, prime contractors with data-residency or US-person requirements, or organizations whose compliance framework requires dedicated, US-owned infrastructure with no shared-tenancy risk. The shared tier (Option A) is explicitly scoped to unclassified / FOUO-and-below data; anything at or above CUI routes to this tier. All infrastructure is managed by IT Pro Partner - the client never touches servers - but the hardware, storage, and backup pipeline are theirs alone.

Shared responsibility line: In both models, ITPP manages everything below the application layer (OS, Docker, database, backups, monitoring, patching, incident response). The client's only responsibility is user account management, profile configuration, and content submitted to the platform. This is the same operational model ITPP uses for all managed infrastructure - the client gets the benefit of bare-metal performance without any of the operational burden.

5.2 The Seven Core Features

Feature 1: Smart Discovery (SETTLED - core design)

Contractor builds a one-time profile: NAICS codes, set-aside eligibility (8(a), SDVOSB, WOSB, HUBZone), geographic preferences, contract size range, agency history. RFP Tank ingests SAM.gov daily and scores every new solicitation against the profile. User receives a morning digest: "3 new opportunities match your profile. 1 is a recompete at an agency where you have past performance. 1 closes in 14 days." Recompete detection flags expiring contracts the user has previously bid or won. When a tracked solicitation is amended, RFP Tank notifies the contractor with exactly what changed - a moved due date, a new clause, an added requirement - not just "an amendment was posted."

Feature 2: RFP Shredder (SETTLED - core design)

Paste a SAM.gov URL, upload a PDF, or drop in a Word document. In 60 seconds: extracts all requirements, maps them to Section L (instructions) and Section M (evaluation criteria), identifies mandatory vs. desirable criteria, flags ambiguous language ("will be evaluated favorably" - what does that mean?), estimates competition level based on set-aside type and historical award data.

Same input path for the contractor's own proposal: upload the response document or paste its URL, and RFP Tank scores it against the RFP. This is the two-input model - RFP in, proposal in, compliance score out. Proposals are often multi-file (technical approach, cost volume, attachments), so ingestion accepts 1..N documents as one submission, not a single file.

Output: a structured compliance matrix. Every requirement, its source section, its evaluation weight, and a response field. This alone saves 5-20 hours per RFP.

Feature 3: Guided Response Flow (OPEN - primary differentiator)

NOT "here is an AI draft, good luck." A step-by-step wizard that walks the contractor through every requirement, section by section. For each requirement: write or paste your response, hit Review, and the AI checks: Does this address the requirement? Does it use government-recognized terminology? Is there quantifiable evidence? Does it map to the evaluation criterion? Green/amber/red per requirement. Progress tracker: "You have addressed 23 of 47 requirements. 3 are partially addressed. 2 are at risk."

This is the product nobody else has built. This is the feature that justifies the price.

Feature 4: Compliance Copilot (OPEN - builds on Feature 3)

Real-time AI assistant that watches as you type. Not post-hoc review - live checking. "This paragraph addresses Requirement 4.2 but does not quantify the outcome. Add a metric (e.g., reduced processing time by 40%)." Tracks every requirement across the full response. Nothing slips through.

Feature 5: Past Performance Vault (OPEN - Team+ tier)

A structured library of every past contract, capability statement, and relevant project description. AI mining: when a new RFP comes in, the vault surfaces relevant past performance automatically. "For this IDIQ requirement around IT modernization, you have 3 relevant past performance references. Click to insert and customize." Over time, the vault becomes the contractor's most valuable asset. It gets smarter with every proposal.

Feature 6: Win/Loss Intelligence (OPEN - Team+ tier)

Track every bid submitted through the platform. Mark wins and losses as awards are announced. Over time, the AI identifies patterns: win rate by agency, by NAICS, by contract size, by proposal section quality score. "Your win rate on DoD contracts is 42%. On DHS contracts it is 67%. Your strongest proposals score 85%+ on the Technical Approach section. Your weakest score below 60% on Management Approach." That insight is worth more than the subscription cost.

Feature 7: One-Click Export (SETTLED - standard)

Export the complete proposal in Word format (formatted per RFP instructions), PDF, or SharePoint-compatible package. Compliance matrix exports to Excel. Auto-generated cover letter. Everything formatted, nothing left to clean up.

5.3 The Differentiation Layer: Traceability + Evaluation

The seven core features are the workflow. This layer is the moat. The gold standard is not a longer feature list - it is depth on the one thing every competitor does shallowly (traceability), plus an evaluation layer none of them do at all.

The three-product handoff keeps lanes clean: RFP Tank owns everything UP TO writing (extract, map, gap-check, score). ProposalTank writes the prose. VerdictTank reviews the finished draft. No feature bleed between them.

A. Traceability (the moat)

B. Evaluation (the differentiator, leverages the VerdictTank panel)

C. Mechanical compliance (boring but disqualifying)

An automated gate for page count, font, margins, per-volume page limits, required forms and signatures, and naming conventions. The top reason proposals are rejected unread.

D. Make-writing-a-breeze (still RFP Tank lane, no prose)

Prioritized build wedge

  1. Compliance matrix + gap report (core deliverable, biggest moat).
  2. Mechanical compliance gate (cheap, high perceived value).
  3. Evaluator simulation (the "will it win" layer).
  4. Real-time side-by-side scoring (the "breeze").
  5. Outline generator + terminology mirroring + amendment re-diff.

5.4 Build vs. Buy Status

Component Status Effort Notes
SAM.gov API integration TO BUILD Low Free API, well-documented, standard REST
USASpending.gov integration TO BUILD Low Free API, historical award data
Opportunity scoring engine TO BUILD Medium Profile-match algorithm, proprietary logic
RFP document parser (PDF/Word) TO BUILD Medium Existing OSS libraries available
Compliance matrix generator TO BUILD Medium-High Core IP, requires FAR/RFP domain logic
Guided response wizard TO BUILD High Primary differentiator, custom build required
Compliance Copilot (real-time AI) TO BUILD High LLM integration + real-time feedback loop
Past performance vault TO BUILD Medium Document storage + vector retrieval
Win/loss analytics TO BUILD Medium Database + dashboard, standard BI patterns
Word/PDF/Excel export TO BUILD Low-Medium Existing OSS libraries available
Authentication + billing TO BUILD Low Auth0 + Stripe, commodity
Marketing site (rfptank.com) TO BUILD Low Standard ITPP deployment

Infrastructure: Runs on existing ITPP app servers (app1/app3). No new servers required at MVP. LLM API costs are variable and passed through via usage-aware tier limits.

5.5 Deployment Status Grid

Site Status Detail
rfptank.com OPEN - not registered Domain confirmed available, needs registration
app.rfptank.com OPEN - not built Main SaaS application
api.rfptank.com OPEN - not built Backend API
rfptank.com (marketing) OPEN - not built Marketing and signup site

5.6 Pricing Tiers

Tier Price Users Core Features
Solo $79/mo ($948/yr) 1 Smart Discovery, RFP Shredder, Guided Response Flow, 5 active RFPs/mo
Consultant $249/mo ($2,988/yr) 1 Everything in Solo + Compliance Copilot, Past Performance Vault (25 entries), Win/Loss tracking, 20 active RFPs/mo
Team $599/mo ($7,188/yr) Up to 5 Everything in Consultant + multi-user collaboration, unlimited Past Performance Vault, advanced Win/Loss analytics, priority support, 50 active RFPs/mo
Enterprise Custom (starts ~$2,000/mo) Unlimited Everything in Team + white-label, API access, dedicated CSM, custom NAICS/agency configurations, SLA, SSO

Pricing rationale: GovEagle does not publish pricing and serves mid-to-large primes with a full sales cycle. We charge $79/month and serve the other 99% - walk-up, self-service, no demo required. The Solo tier is priced to convert on impulse. The Consultant tier is priced where a single proposal win pays for 5 years of the subscription. The Team tier is priced transparently at $599/month for up to five users, with no demo or sales call required.


6. Revenue Model

6.1 Revenue Streams

Stream Source Year 1 Revenue Notes
Solo subscriptions $79/mo per user $148,520 157 subscribers avg
Consultant subscriptions $249/mo per user $185,505 62 subscribers avg
Team subscriptions $599/mo per account $138,369 19 accounts avg
Total Year 1 Revenue $472,394 Conservative ramp (619 accounts at Month 12)
End-of-year ARR run rate $1,278,972 Month 12 MRR $106,581 x 12

Enterprise contracts (custom, ~$2,000/mo) are excluded from the base model; 2-3 added in Q4 would push Year 1 total above $500K.

6.2 Pricing Rationale - Why These Numbers Are Right

This is not budget pricing. It is premium pricing positioned below enterprise alternatives. Here is the ROI math for each tier:

Solo at $79/month:

A solo contractor pursuing 10 RFPs per year spends 40-100 hours per response. At a blended rate of $100/hour in opportunity cost, that is $40,000-$100,000 in annual labor. Saving 30% of that is $12,000-$30,000 in value. The subscription costs $948/year. ROI is 12:1 to 30:1. This is a no-brainer at double the price. We price it at $79 to make the conversion decision require zero deliberation.

Consultant at $249/month:

A GovCon consultant who wins a single $300K task order earns a fee of $30K-$60K. If RFP Tank helps them win one additional contract per year that they would have otherwise lost, the subscription ($2,988/year) is irrelevant noise against a $30K fee. This tier also unlocks the Past Performance Vault - which becomes the consultant's competitive moat over time. Priced at $249 for transparent self-service access to an enterprise-grade feature set that GovEagle gates behind an unpublished enterprise sales process.

Team at $599/month:

A 5-person BD team at a small GovCon firm costs $300K-$600K in annual salaries. A 10% improvement in proposal quality that lifts win rate from 30% to 33% on a $10M pipeline generates $300K in additional contract awards. The subscription costs $7,188/year. The comparison is not close. We price at $599 because that is where a decision-maker can approve it without a committee.

Enterprise at Custom:

Enterprise accounts start at approximately $2,000/month and are scoped based on user count, API access, white-label requirements, and SLA needs. This is where the GovEagle comparison lands - we offer a comparable enterprise feature set with transparent pricing and self-service onboarding, versus GovEagle's unpublished enterprise pricing and 1-week implementation requirement.

6.3 Cost Structure

Cost Item Monthly Annual Notes
LLM API (OpenAI/Anthropic) $800-$2,500 $9,600-$30,000 Scales with usage; tier limits control cost
Server infrastructure $0 $0 Already running on ITPP infra
SAM.gov API $0 $0 Free public API
USASpending.gov API $0 $0 Free public API
Auth0 (authentication) $70 $840 Up to 7,000 MAU on free tier initially
Stripe (payment processing) 2.9% + $0.30/transaction ~$2,000-$8,000 Variable with revenue
Domain + SSL $15 $180 rfptank.com registration
Monitoring + logging $50 $600 Datadog or equivalent
Customer support tooling $100 $1,200 Help desk software
Marketing (content + ads) $2,000 $24,000 Primarily content + LinkedIn
Total Monthly OpEx ~$3,000-$5,000 (early months) ~$36,000-$60,000 Ramps to ~$11,000/month at scale (see 10.1)

Key insight: The infrastructure advantage is real. Other startups in this space are paying $5,000-$15,000/month in AWS/GCP costs before they write a line of product code. ITPP infrastructure is already running. That is a structural cost advantage that compounds as we scale.

6.4 Gross Margin

At full scale (619+ accounts), estimated gross margin is 78-85%. The primary variable cost is LLM API usage, which is manageable through tier-based usage limits (RFPs per month, compliance checks per day). A single compliance check is 1-2 LLM calls of roughly 1,000-2,000 tokens, or about $0.02-$0.05 at current Claude/GPT pricing; a heavy session of 50 checks costs $1.00-$2.50, leaving the $79/month Solo tier at 80%+ gross margin even under daily-use limits. Real-time copilot latency is bounded to roughly 1-2 seconds per check with streaming, acceptable for an inline writing assistant. Unlike a services business, the marginal cost of serving the 501st customer is near zero.

6.5 12-Month Revenue Ramp

Month Solo Subs Consultant Subs Team Accounts MRR Cumulative Revenue
1 10 2 0 $1,288 $1,288
2 20 5 1 $3,424 $4,712
3 35 10 2 $6,453 $11,165
4 55 18 4 $11,223 $22,388
5 80 28 7 $17,485 $39,873
6 110 40 11 $25,239 $65,112
7 145 55 16 $34,734 $99,846
8 185 72 22 $45,721 $145,567
9 230 92 29 $58,449 $204,016
10 280 115 37 $72,918 $276,934
11 335 140 46 $88,879 $365,813
12 395 168 56 $106,581 $472,394

Year 1 total revenue: ~$472K (conservative ramp, 0% churn assumption in early months is unrealistic - see Risk section).

End-of-year MRR run rate: $106,581 ($1.28M ARR) across 619 paying accounts.

These numbers assume no enterprise contracts in Year 1. With 2-3 enterprise accounts added in Q4, Year 1 total exceeds $500K and end-of-year ARR exceeds $1.3M.


7. Competitive Advantages (Moat)

7.1 The Seven Knockout Blows

These are not features. They are structural advantages that compound over time. Each one is hard to copy individually. Together, they are a fortress.

To be precise about what is and is not durable: three of the seven are structural advantages that take real build time and GovCon domain knowledge (profile-driven matching, guided response flow, real-time compliance). Two are positioning choices a competitor could copy in a quarter if it decided to (self-service pricing, full-stack coverage). The durable, compounding moat is Knockout 4 plus Section 7.3 - the Past Performance Vault and win/loss data lock-in, which cannot be copied because the value lives in the user's own data, not in our code.

Knockout 1: Profile-Driven Proactive Matching

Every other tool is reactive - you go find the RFP. RFP Tank finds the RFP for you, scores it against your profile, and tells you which ones are worth your time before you ever open the solicitation. Nobody else does this end-to-end. It requires building a profile engine, a scoring algorithm, and a daily ingestion pipeline. CLEATUS does partial matching but without the full profile context or the scoring depth.

Knockout 2: Guided Response Flow

This is the product nobody built. It is the feature that should exist and does not. The reason: building it correctly requires understanding how government proposals are evaluated - Section L/M alignment, FAR clause awareness, evaluation factor weighting. That knowledge is not in engineering departments. It is in proposal professionals. We are building it first, ahead of everything else, because it is the differentiator the rest of the product hangs on. No competitor has it, and none is positioned to build it quickly.

Knockout 3: Real-Time Compliance Copilot

Every other AI in this space does post-hoc review: "Here is a score for your completed proposal." We check in real time, as you write. The distinction is not cosmetic - catching a compliance gap while writing is 10x cheaper than catching it after you have written 40 pages that need to be rewritten. This is a workflow advantage that shows up in win rates, not just user experience.

Knockout 4: Past Performance Vault With AI Mining

Your past performance is your biggest competitive asset in government contracting. It is also scattered across email chains, old Word documents, and the memory of your BD team. The Vault organizes it, and the AI mines it automatically for each new opportunity. Over time, users who have populated their vault have a compounding advantage over users who have not - and over competitors whose tools do not offer this capability at all.

Knockout 5: Win/Loss Analytics

This is CRM for proposal professionals. Track every bid, mark outcomes, and let the AI identify what separates your wins from your losses. No competitor offers this at the SMB level. Enterprise BD teams at large primes do this manually with spreadsheets. We automate it for solo contractors and 5-person BD teams at a fraction of the cost.

Knockout 6: Transparent Self-Service Pricing

The GovCon software market is dominated by "book a demo" gatekeeping. GovEagle, GovDash, Vultron, Rohirrim - none of them will tell you what they charge without a sales call. That is a conversion killer for small contractors who do not have time for a 3-meeting sales cycle. RFP Tank has a pricing page. You can be live in 5 minutes. No demo. No call. No contract negotiation. That is a product-led growth strategy that scales without a sales team.

Knockout 7: Full-Stack Workflow for Any Contractor Size

Every competitor serves a slice of the market. GovEagle serves large primes. CLEATUS serves individuals (with a shallow product). Nobody serves the full spectrum - from solo consultant to 50-person firm - with a single product that scales appropriately. RFP Tank does. One platform, four tiers, zero workflow gaps. A solo contractor who grows to a 10-person firm does not need to switch tools. They upgrade their tier.

7.2 Competitive Positioning Map

                HIGH CAPABILITY
                      |
          RFP Tank    |   GovEagle
          [Solo-Ent]  |   [Enterprise]
                      |
SELF-SERVICE ---------+--------- ENTERPRISE-ONLY
    $79/mo            |              Unpublished
                      |
          CLEATUS     |   GovDash
          [Search]    |   [Price-led]
                      |
                LOW CAPABILITY

RFP Tank occupies the upper-left quadrant - high capability, self-service access. No competitor is positioned there. That is not an accident. Building high capability AND self-service simultaneously is hard. It requires product discipline that most enterprise-focused companies cannot apply because their incentive is to serve enterprise customers with white-glove support, not to make the product simple enough for a solo contractor to onboard in 5 minutes.

7.3 The Moat Gets Deeper Over Time

The Past Performance Vault creates data lock-in. Users who have spent 6 months populating their vault with past contracts, capability statements, and proposal content are not switching tools. The switching cost is not financial - it is the loss of an organized, AI-indexed library that took months to build. This is the same moat Salesforce uses (CRM data lock-in), applied to the GovCon proposal workflow.

Win/Loss data deepens the moat further. After 12 months of tracking bids and outcomes, the platform knows which proposal sections correlate with wins for that specific contractor, at that specific agency, in that specific NAICS. That intelligence is contractor-specific and cannot be transferred to a competitor's tool.


8. Go-to-Market Strategy

8.1 Phase Structure

Phase Timeline Goal Key Activities
Foundation Months 1-2 Build the product + pre-launch waitlist Core feature development, rfptank.com marketing site live, LinkedIn presence, waitlist at 200+
Soft Launch Month 3 First 50 paying customers Invite waitlist, APEX Accelerator outreach, GovCon LinkedIn content, beta feedback loop
Growth Months 4-8 Reach 200 paying customers Content marketing, GovCon forum presence, referral program, first case studies
Scale Months 9-12 500+ subscribers, first enterprise accounts Paid LinkedIn ads, APMP partnerships, conference presence, enterprise sales motion

8.2 First 100 Customers - Where They Come From

  1. LinkedIn GovCon community - There is a large, active community of government contractors on LinkedIn. Posts about proposal pain points, compliance matrices, and SAM.gov go viral in this community. A single high-quality post from a credible voice ("I built a tool that shreds an RFP into a compliance matrix in 60 seconds") reaches thousands of relevant people organically.
  1. APEX Accelerator offices - APEX Accelerators (formerly Procurement Technical Assistance Centers) exist in every state. They are funded by the DoD to help small businesses win government contracts. They actively look for tools to recommend to their clients. A 30-minute conversation with an APEX Accelerator director can unlock referrals to 50-200 contractors in that region.
  1. 8(a) cohort lists - The SBA publishes lists of 8(a)-certified firms. These are small disadvantaged businesses in the program specifically to win government contracts. They are exactly our user. LinkedIn outreach to 8(a) firm principals converts at high rates because their pain is acute and our solution is direct.
  1. GovCon forums and communities - GovLoop, Govcon Giant (Joshua Frank's community), NCMA (National Contract Management Association) members, APMP (Association of Proposal Management Professionals). These communities have high concentrations of our exact target user.
  1. APMP partnership - The Association of Proposal Management Professionals is the professional association for proposal writers. An endorsement or partner listing here reaches the most sophisticated buyers in the space and lends immediate credibility.
  1. Germaine's existing network - IT Pro Partner has existing relationships with government contractors, MSP clients, and the broader business community. Warm introductions from a trusted advisor convert at 3-5x the rate of cold outreach.
  1. Veteran business networks - SDVOSB (Service-Disabled Veteran-Owned Small Business) owners are a concentrated, self-identified group with acute proposal pain. VetBiz, American Legion posts, SCORE chapters with veteran programs.
  1. Content marketing - "How to build a compliance matrix in 30 minutes" is a top-searched query among GovCon professionals. A library of practical how-to content on rfptank.com drives organic search traffic from exactly the right audience.
  1. GovWin / Deltek community - GovWin (now owned by Deltek) charges $40K/year and has an enormous dissatisfied customer base. Anyone who has ever searched for a GovWin alternative is a perfect target. SEO and LinkedIn ads targeting GovWin users convert efficiently.
  1. Twitter/X GovCon accounts - A handful of high-follower GovCon educators and consultants on X have audiences of 10K-50K exactly aligned with our target user. A single mention or repost from one of these accounts can drive hundreds of signups.

8.3 Channel Strategy

Channel Tactic Cost Expected Reach Expected Conversion
LinkedIn organic Weekly GovCon content, founder posts about the product $0 2,000-10,000 per post 1-3% to signup
APEX Accelerator partnerships Email outreach to 50 regional APEX Accelerators $0 50-200 referrals per active APEX Accelerator High - warm referrals
LinkedIn paid ads Target: NAICS code keywords, 8(a) designation, SAM.gov users $2,000-$5,000/mo 50,000-100,000 impressions 0.5-1% CTR, 10% trial conversion
Content / SEO RFP compliance, proposal writing how-to articles $500/mo (writer) Compound - 5,000-20,000 visits/mo by Month 12 2-4% trial signup
GovCon forums Authentic community participation + tool mentions $0 500-2,000 per active thread 5-10% (trust-based)
APMP membership Attend chapter meetings, sponsor newsletter $1,000-$3,000/yr 5,000 proposal professionals 2-5% trial signup
Referral program 2 months free for each paying referral Revenue share Compounds with customer base 30-50% of growth by Month 9
Cold outreach LinkedIn DM to 8(a) and SDVOSB firm principals $0 100 contacts/week 5-10% response, 2-3% trial

All conversion rates above are planning hypotheses based on benchmark SaaS funnels and adjacent GovCon communities; none have been field-tested. The first 90 days are scoped to validate these against real channel data before scaling spend (see 9.6).

8.4 Product-Led Growth Hook

The RFP Shredder is a natural free-trial hook. Let any user - without an account - paste a SAM.gov opportunity link and get a preview compliance matrix (first 10 requirements extracted, rest gated). This creates an immediate "holy sh*t" moment. The prospect sees the value before they give us an email address. When they hit the gate on requirement #11, they sign up. This is how Figma, Loom, and Notion grew their initial user bases - lead with the product, not the pitch.


9. Risk Analysis

9.1 Market Risks

Risk Probability Impact Mitigation
Federal contract spending cuts (budget sequestration) Low High SLED market ($1.6T) is not federally funded; pivot emphasis to state/local if federal spend contracts significantly
Government moves RFP process fully to AI-generated solicitations, reducing proposal complexity Very Low High This is 10-15 years away at minimum; FAR reform is glacial
SAM.gov API changes or rate limiting Medium Medium Build abstraction layer; maintain cached opportunity database; FedBizOpps has multiple unofficial aggregators as backup
GovCon market contraction (agency consolidations) Low Medium More consolidation typically means larger contracts with more competitive bidding - a tailwind for proposal tools
Small business set-aside percentages reduced by policy change Very Low Medium Total contract spend growing regardless; larger share of smaller pie still represents accessible market

9.2 Product Risks

Risk Probability Impact Mitigation
Compliance matrix AI produces errors that cause proposal disqualification Medium High Build in explicit disclaimer; require user review of all AI output; red-flag high-stakes requirements for manual verification; never position AI as a replacement for expert review
LLM API costs exceed projections as usage scales Medium Medium Implement strict per-tier usage limits from Day 1; build caching for common document types; negotiate enterprise API pricing above 10M tokens/month
60-second RFP shredder cannot handle all document formats (scanned PDFs, unusual formatting) High Low-Medium Build graceful degradation - partial extraction is better than failure; queue for manual review on complex documents
Guided response flow UX is too complex for solo contractors Medium High Conduct user testing with 10-20 target users before launch; iterate on flow based on completion rates; provide skip/simplify options
Past performance vault has slow adoption (users do not populate it) High Low Build import tools (upload old proposals, pull from USASpending); seed with public data; gamify completion

9.3 Competitive Risks

Risk Probability Impact Mitigation
GovEagle launches a self-service tier Low High Their sales model, org structure, and positioning make this nearly impossible to execute without a complete company pivot; we would still have 12-18 months of market lead
CLEATUS adds a guided response flow Medium Medium CLEATUS is a data company - adding a workflow layer to a search tool requires rebuilding the product; their feature velocity is slow
New entrant raises $10M+ to build exactly what we are building Medium High Speed is the mitigation; launch in 90 days, get to 100 customers before they get to product-market fit; first-mover + data moat compounds quickly
OpenAI or similar releases a general-purpose "RFP assistant" Medium Low-Medium Generic tools cannot replicate the SAM.gov integration, profile matching, and FAR-specific compliance logic; government contracting is too specialized for a general assistant
GovDash copies the guided response flow Low Medium Their positioning is "cheap alternative" - adding a premium workflow feature contradicts their own marketing

9.4 Operational Risks

Risk Probability Impact Mitigation
Single-developer bottleneck (Germaine) High High Document architecture thoroughly from Day 1; pre-commit a contract developer for the SAM.gov ingestion pipeline and compliance-matrix backend from Week 4 (fixed ~$2,000-$4,000, not gated on traction); open source non-core components to attract contributors
Customer support load overwhelms available capacity Medium Medium Invest in self-service documentation from launch; build FAQ and tutorial library before the first 50 customers; implement ticket triage by tier
Data breach or security incident exposes contractor proposal data Low Very High Contractor proposals contain competitively sensitive pricing, teaming information, and proprietary technical approaches; encrypt all data at rest and in transit; implement SOC 2 Type II audit by Year 2; no CUI handling on the shared tier (FedRAMP not required there); CUI and ITAR workloads route to the US-owned dedicated tier (Option B), built to DFARS 252.204-7012 / FedRAMP Moderate-equivalent standards
Key customer (enterprise) churns in first 6 months Medium Medium Enterprise customer success process from Day 1; dedicated Slack channel for enterprise accounts; monthly check-ins; usage monitoring to catch disengagement early
No assumption validated with a customer yet High High Every number in this document (willingness-to-pay, conversion rates, CAC) is a hypothesis until tested; run 10 discovery interviews and a waitlist before scaling marketing spend; kill or reprice tiers the data does not support

9.5 Pre-Mortem: What Kills This Within 6 Months

The most likely failure modes, in order of probability:

1. The guided response flow ships too slow. If the differentiating feature - the guided wizard with real-time compliance checking - takes 6 months to build, we launch with just the RFP Shredder. The Shredder is useful, but it is not a sufficient moat. GovDash and GovEagle both have compliance matrix generation. Without the guided flow, we are just another RFP tool. Mitigation: the guided flow is feature #1, not feature #7. Build it first, ship everything else after.

2. Trial-to-paid conversion is below 5%. If users love the free RFP Shredder preview but do not convert to paying subscribers, the product-led growth loop breaks. This usually means the paywall hits before users see enough value, or the paid features are not differentiated enough from the preview. Mitigation: track conversion rates weekly from Day 1; iterate the paywall and trial experience aggressively in the first 60 days.

3. LLM API costs eat the margin. At scale, if users are running dozens of compliance checks per session and the per-check LLM cost is not bounded, gross margin compresses to zero. Mitigation: enforce per-tier daily limits from Day 1; cache common document parsing results; monitor API spend daily.

4. GovCon community does not trust a new tool with proposal data. Government contractors are paranoid about data security for good reason - proposals contain pricing strategies, teaming agreements, and proprietary technical approaches. If RFP Tank cannot answer basic security questions (encryption, data retention, who can access my data), the trust barrier kills conversion. Mitigation: publish a clear security page before launch; answer these questions proactively on the pricing page.

5. Product-market fit is in Consultant tier, not Solo. The economics of Solo ($79/month) require high volume to generate meaningful revenue. If product-market fit turns out to live in the Consultant ($249) and Team ($599) tiers - which is plausible, since those users have more acute pain and more financial stake in win rates - the marketing strategy needs to tilt toward that segment immediately. Mitigation: track ARPU and NPS by tier; respond to where the enthusiasts actually are.

9.6 90-Day Success Metrics

Metric Target Failure Signal
Waitlist signups before launch 500+ Under 200 = insufficient demand signal
Trial signups in first 30 days post-launch 100+ Under 50 = distribution problem
Trial-to-paid conversion 20%+ Under 10% = pricing or value proposition problem
Month 3 paying customers 50+ Under 25 = growth rate will not hit Year 1 target
Net Promoter Score (NPS) 40+ Under 25 = product is not delivering on promise
Churn in first 90 days Under 9% Over 15% = product-market fit problem
Average session length 20+ minutes Under 10 minutes = users are not engaging with the workflow

10. Financial Projections

10.1 12-Month P&L

Month Revenue LLM API Marketing Labor (est.) Other OpEx Net Income
1 $1,288 $200 $1,000 $5,000 $500 -$5,412
2 $3,424 $400 $1,500 $5,000 $500 -$3,976
3 $6,453 $600 $2,000 $5,000 $600 -$1,747
4 $11,223 $900 $2,500 $5,000 $700 $2,123
5 $17,485 $1,200 $3,000 $5,000 $800 $7,485
6 $25,239 $1,600 $3,500 $5,000 $900 $14,239
7 $34,734 $2,000 $4,000 $5,000 $1,000 $22,734
8 $45,721 $2,400 $4,500 $5,000 $1,000 $32,821
9 $58,449 $3,000 $5,000 $7,500 $1,200 $41,749
10 $72,918 $3,500 $5,000 $7,500 $1,200 $55,718
11 $88,879 $4,000 $5,000 $7,500 $1,500 $70,879
12 $106,581 $4,500 $5,000 $7,500 $1,500 $88,081
TOTAL $472,394 $24,300 $42,000 $70,000 $11,400 $324,694

Notes on assumptions:

10.2 Unit Economics

Metric Solo Consultant Team Enterprise
Monthly price $79 $249 $599 ~$2,000
Gross margin (estimated) 82% 80% 78% 75%
Assumed monthly churn 3% 2% 1.5% 1%
Average customer lifetime 33 months 50 months 67 months 100 months
LTV (lifetime value, gross profit) $2,138 $9,960 $31,304 $150,000
Estimated CAC (all channels) $80 $150 $400 $2,000
LTV/CAC ratio 27:1 66:1 78:1 75:1

LTV/CAC ratios above 3:1 are considered healthy for SaaS. Ratios above 10:1 indicate potential underinvestment in growth. All four tiers show ratios in the 27:1 to 78:1 range. Note these ratios rest on CAC figures that are planning estimates, not field-tested channel data; they indicate room to increase acquisition spend once real CAC is measured, not permission to spend ahead of validation. On churn: the 3% monthly Solo assumption compounds to ~8.7% over 90 days, just under the 'Under 9%' success threshold in section 9.6 - a deliberately tight margin that makes early churn the model's most sensitive assumption and the primary Month-4 go/no-go gate.

10.3 Break-Even Analysis

At the early-stage cost structure from 6.3 (~$3,000-$5,000/month in OpEx excluding labor), the business turns cash-flow positive in Month 2 and net-income positive (including a $5,000/month labor allocation for Germaine's opportunity cost) in Month 4, at 77 paying subscribers and ~$11,200 MRR.

The $10,000-$12,000/month OpEx figure is the Month 11-12 scale structure (see 10.1), reached only after the business is running ~$90,000/month in net income. Break-even is an early-months event, not a scale-stage milestone - a realistic and achievable timeline for a product with genuine product-market fit.

10.4 Three-Scenario Model (Year 1)

Scenario Subscriber Growth Year 1 Revenue End ARR
Conservative 15%/mo from Month 3 $472K $1.28M
Realistic 20%/mo from Month 3 + 3 enterprise $520K $1.45M
Aggressive 30%/mo from Month 3 + 6 enterprise, partnership channels activate $850K $2.1M

The conservative scenario requires 619 paying accounts by Month 12. The realistic scenario adds enterprise contracts and assumes one major APEX Accelerator or APMP channel activation. The aggressive scenario assumes a product that becomes the default recommendation in GovCon communities by Month 6 - plausible if the guided response flow delivers on its promise.

10.5 3-Year Outlook

Year Subscribers ARR Gross Profit Notes
Year 1 619 $1.28M $1.0M Build + launch + first traction
Year 2 2,000 $5.1M $4.0M Scale GTM, hire sales + support
Year 3 6,000 $16.2M $13.0M Enterprise motion + market leadership

Year 3 ARR of $16M at 80% gross margin is a $130M+ valuation at standard SaaS multiples (8-10x ARR). This is not a lifestyle business. It is a venture-scale outcome from a $0 infrastructure investment and a domain that costs $15/year.


11. The Ask

11.1 Resources Needed to Launch

Item Details Timeline Estimated Cost
Domain registration (rfptank.com) Confirmed available - register immediately Day 1 $15/year
SAM.gov API key Free public registration at api.sam.gov Week 1 $0
USASpending.gov API key Free public registration Week 1 $0
LLM API access (OpenAI or Anthropic) Production API key with usage monitoring Week 1 $50-$200 seed deposit
Stripe account setup Payment processing for subscriptions Week 1 $0 upfront
Auth0 account setup Authentication (free tier to 7,000 MAU) Week 1 $0 initially
Marketing site (rfptank.com) Deploy on existing ITPP infrastructure Weeks 1-2 $0
Customer discovery interviews 10 GovCon contractor interviews validating willingness-to-pay and feature priorities Weeks 1-3 $0
Core product development SAM.gov ingestion + RFP Shredder + Guided Flow MVP Months 1-2 Germaine's time
Pre-launch LinkedIn content 8-10 posts building waitlist Months 1-2 $0
APEX Accelerator outreach campaign Email to 20-30 regional APEX Accelerator offices Month 2 $0
Beta program (10-20 users) Invite from waitlist for pre-launch feedback Month 3 $0 or comped subscriptions
Contract development support SAM.gov ingestion + compliance-matrix backend Weeks 4-10 $2,000-$4,000
Launch marketing push LinkedIn, content, community announcements Month 3 $500-$1,000

Total hard dollar cost to launch: ~$2,500-$5,000 (contract-developer contingency plus marketing), removing the single point of failure without waiting for traction. The rest is ITPP infrastructure (already paid for) and Germaine's time. This is the single best risk/reward ratio in the current ITPP product portfolio.

11.2 Immediate Decisions Required

These decisions need answers before Week 2. Each one unblocks a subsequent action.

Decision 1: Domain - Register rfptank.com NOW

Domain Status Verdict
rfptank.com AVAILABLE - confirmed Top pick - direct, memorable, product-literal. Register today.
rfptank.io Check availability Acceptable fallback if .com somehow expires before registration
rfpshredder.com Not checked Describes one feature, not the full product - weaker brand
govproposal.com Not checked Generic, not memorable
proposaltank.com Not checked Acceptable alternative if rfptank.com ever became unavailable

Recommendation: Register rfptank.com today. Domains at this price point that are available do not stay available if the concept leaks. This is a $15 decision with potentially significant downstream value. There is no reason to delay.

Decision 2: Product vs. Consulting Play

RFP Tank is an IT Pro Partner SaaS product. It is not a GovCon consulting practice. The product helps others win contracts; ITPP does not bid contracts using it (yet). There is an optional future play where Germaine's veteran colleague and ITPP form a minority-veteran-owned JV that USES RFP Tank to bid on IT contracts - creating live case studies and a demonstration environment. But the product itself launches as ITPP software. Confirm: product play first, consulting JV optionally later.

Decision 3: Build Sequence

The guided response flow is the differentiator. It should be Feature 1, not Feature 7. Proposed sequence:

  1. SAM.gov ingestion + profile matching (2 weeks)
  2. RFP Shredder + compliance matrix (2 weeks)
  3. Guided response wizard MVP (4 weeks)
  4. Marketing site + billing + auth (2 weeks)
  5. Beta launch (Week 12)

This sequence carries two weeks of built-in slack. The guided response wizard MVP (4 weeks) is the critical path, and the daily SAM.gov ingestion is contracted out (see 11.1) so it does not compete for Germaine's time. If the guided flow slips, beta still ships on the RFP Shredder + profile matching (already differentiated), and the guided flow lands in Week 12.

Confirm this build order before Day 1 of development.

Decision 4: Pricing Confirmation

The tiers proposed ($79/$249/$599/custom) are recommendations. They are priced to be premium without being enterprise. If Germaine wants to price higher (e.g., Solo at $99, Consultant at $299) the unit economics only improve. If the goal is maximum early traction, the current tiers are correct. Confirm pricing before the marketing site goes live.

Decision 5: IT Pro Partner Branding

Does rfptank.com show "Powered by IT Pro Partner" or does it stand alone as a brand? Most SaaS products benefit from standalone branding for credibility in a niche market. ITPP is an MSP brand; RFP Tank is a GovCon software brand. They serve different audiences. Recommendation: standalone brand for rfptank.com, with a discreet "A product of IT Pro Partner" attribution in the footer.

11.3 What Success Looks Like at Month 12

Metric Target
Paying accounts 619+
Monthly Recurring Revenue $106,000+
Annual Recurring Revenue $1.28M+
Net Promoter Score 45+
Churn rate (monthly) Under 3%
Trial-to-paid conversion 20%+
Enterprise accounts 2-5
APEX Accelerator partnerships active 5+ regional APEX Accelerators recommending RFP Tank
Case studies published 3+ contractors with documented win-rate improvements

The benchmark for "product-market fit confirmed" is Month 4: 50+ paying subscribers who were NOT given the product for free, with NPS above 40 and monthly churn below 3%. If those three numbers are green at Month 4, this product grows to $1.28M ARR. If they are red, we know what to fix before we have spent real marketing dollars.

11.4 The Bigger Picture

RFP Tank is not just a product. It is a position in a $793 billion market that nobody has credibly claimed at the self-service level. The government contracting industry is underserved at the small business layer because the tools that exist were built for the large primes, not for the contractors who need them most.

Every month that passes without RFP Tank is a month that 50,000+ small business contractors submit proposals in Word, built with spreadsheet compliance matrices, without a single line of AI assistance on whether they actually addressed the requirement they think they addressed. Some of them will lose contracts they deserved to win. Some of them will exit the GovCon market entirely, convinced the system is rigged, when the real problem was a solvable workflow failure.

RFP Tank does not just serve the market. It corrects a market failure at scale. That is the kind of business worth building.

The competition saw what was possible in this market. They served the 1% and ignored the rest. That was their plan. And then RFP Tank is going to hit them in the mouth.


Appendix A: Competitor Pricing Deep Dive

Competitor Public Pricing Actual Pricing (Intel) Pricing Model Sales Motion
GovEagle None published Unpublished (enterprise) Enterprise custom Demo required, 1-week onboarding
GovDash None published Custom, modular Seat-free, module-based Demo required
CLEATUS Published $39/mo (DATA), $78/mo (DATA+AI) Self-service Walk-up
GovSignals None published Custom enterprise Demo required Sales-led
Vultron None published Custom (raised $22M Series A) Demo required Sales-led
Rohirrim None published Custom enterprise Demo required Sales-led
GovWin (Deltek) None published ~$40,000/year Enterprise Full sales cycle
GovTribe None published ~$25,000/year Enterprise Full sales cycle

Observation: Only CLEATUS publishes walk-up pricing in the self-service segment. Every other player gates pricing behind a sales call. RFP Tank and its published tier structure will be immediately differentiated from every enterprise competitor the moment the pricing page goes live.

Appendix B: Data Source Summary

Data Source Cost Data Available API Quality
SAM.gov Free (API key required) All federal solicitations, contract awards, entity registrations Well-documented, REST, reliable
USASpending.gov Free Award data, historical pricing, agency spend by NAICS Good, bulk download available
FPDS (Federal Procurement Data System) Free Contract awards, modifications Older API, some quirks
State procurement portals Free (scraping) Variable by state No standard API; scraping required for SLED coverage
Federal Register Free Regulatory notices, advance procurement notices XML feeds available

Total data infrastructure cost: $0. The public data is free. The value is not in the data - it is in what you do with it.


Proposal prepared by IT Pro Partner | Germaine Brown | August 2026

RFP Tank is an IT Pro Partner product. rfptank.com is confirmed available as of proposal date.

All market data sourced from SAM.gov, USASpending.gov, SBA.gov, and publicly available competitor research.

2. Subscriber Requirements

Source of truth for subscriber-facing behavior. Captured live during ideation; referenced by the proposal and architecture.

Captured live during ideation with Germaine. One entry per numbered

requirement as they are stated. This file is the source of truth for the

subscriber-facing behavior; the proposal (rfptank-proposal.md) and architecture

(v3.6-architecture.md) reference it.

Requirement 1 - Subscriber proposal ingestion

A logged-in subscriber can submit their proposal for compliance review by either:

Design notes (to resolve before build)

  1. Two ingestion paths, one normalized output. File upload and URL fetch are

different code paths that must both converge on extracted, clean text before

compliance scoring runs. File upload = parse PDF/DOCX/etc. URL = fetch +

extract (PDF download, Google Docs/SharePoint export, or plain page).

  1. This is the second half of the two-input model. The existing proposal

already specs the RFP side ("Paste a SAM.gov URL, upload a PDF, or drop in a

Word document"). Requirement 1 is the PROPOSAL side. RFP Tank scores the

proposal against the RFP, so a full run needs BOTH ingested.

  1. Proposals are often multi-file. Technical approach, cost/price volume,

attachments, past performance. Ingestion must accept 1..N documents and treat

them as one submission, not a single file.

  1. URL path has a hard case. Auth-gated links (Google Docs, SharePoint,

portals behind login) and non-extractable pages are the common failure. The

URL flow needs a clear "can't reach this link" error with a fallback to

manual upload, not a silent empty ingest.

  1. Login gates everything downstream. "Log into the portal" is the first

step of the journey, so auth (shared Stack Auth/Clerk tenant) has to exist

before any ingestion is usable. This is a dependency on the shared auth layer

already decided for ProposalTank + VerdictTank + RFP Tank.

Requirement 2 - Profile NAICS + RFP change notifications

Two parts, both from Germaine's second requirement message.

2a. NAICS codes on the subscriber profile

How many NAICS codes a subscriber can attach to their profile. Purpose: NAICS

is the primary filter for RFP matching/tracking (which opportunities are

relevant to this subscriber).

~99% under 10. Keeps matching signal clean.

unlimited.

business description and auto-suggest related codes from the NAICS hierarchy.

2b. RFP change/amendment notifications for tracked solicitations

When a subscriber is tracking an RFP and that RFP is amended or changed, send

them a notification. Table-stakes feature (all competitors do amendment

alerts; a missed amendment = late or non-compliant proposal).

SAM.gov Opportunities API for that notice on a schedule (every 6-12 hrs),

diffs against last-seen version, notifies on change.

set-aside changed, requirements/description updated.

as a premium add.

requirement), not just "an amendment was posted."

Gold-standard feature brainstorm (2026-08-28)

Frame: the gold standard is not a longer feature list, it is depth on ONE

thing competitors do shallowly - traceability - plus an evaluation layer they

do not do at all. RFP Tank owns everything UP TO writing; ProposalTank writes

the prose; VerdictTank reviews the draft. No feature bleed.

A. Traceability (the moat)

mapped to the exact proposal paragraph that answers it, with confidence and

quoted evidence. Exportable to XLSX/PDF. This is the deliverable a pursuit

manager prints and hands to leadership.

thin coverage = weakness; separate the 3 real risks from the 11 fine items.

wording (evaluators keyword-scan) and where it is vague.

B. Evaluation (the differentiator, leverages VerdictTank panel)

RFP's own evaluation criteria and weights. Answers "will it win", not just

"does it comply". Uses the inherited VerdictTank panel engine.

addressed, each with evidence.

C. Mechanical compliance (boring but disqualifying)

signatures present, naming conventions. Automated gate. Top reason

proposals are rejected unread.

D. Make-writing-a-breeze (still RFP Tank lane, no prose)

each requirement pre-mapped to a section and a page budget per section.

Writer fills in, never starts blank.

updates as they type.

requirements and answers are now stale (extends Req 2b).

Prioritized wedge (build order)

  1. Compliance matrix + gap report (the core deliverable, biggest moat).
  2. Mechanical compliance gate (cheap, high perceived value).
  3. Evaluator simulation (the "will it win" layer).
  4. Real-time side-by-side scoring (the "breeze").
  5. Outline generator + terminology mirroring + amendment re-diff.

3. Baseline Scoring (R3)

The proposal itself run through the RFP Tank compliance-scoring engine - a 10-dimension verdict with cited gaps. Demonstrates the evaluation layer in action. Baseline 4.0/10 was issued on the pre-fix draft; every arithmetic and consistency finding cited below is now corrected in the document above: TAM $10B error -> $480M TAM / $9.6M capture target, GovEagle $30K/yr -> unpublished, Solo LTV $2,607 -> $2,138, SAM 502K/302K -> 200K, churn target Under 5% -> Under 9%, launch cost Under $500 -> $2,500-$5,000, Month-12 MRR $85,072 -> $106,581, PTAC -> APEX Accelerators. The only remaining open item is customer validation (10 GovCon interviews), scheduled for Weeks 1-3 in the roadmap.
4.0
Avg score / 10
10
Dimensions
3
Idea-scored
7
Proposal-scored
DimensionScoreType
Problem Definition7idea
Market Analysis3idea
Competitive Moat3idea
Business Model4proposal
Unit Economics3proposal
Technical Architecture4proposal
Go-to-Market Strategy5proposal
Risk Assessment6proposal
Execution Feasibility3proposal
Growth Trajectory2proposal
Problem Definition 7
The pain is real and specifically drawn: 50-200 page solicitations, manual compliance matrices, Section L/M mapping, and disqualification for non-conformance are genuine GovCon failure modes, and the Section 3.1 method/cost table is the sharpest piece of reasoning in the document. But every number describing that pain is asserted rather than sourced, and the three personas in Section 3.2 are authored composites, not observed users. The problem is plausible; it has not been demonstrated.
Cited gap: Section 3.2 states of Persona 1 'What they will pay: $79/month without hesitation' and of Persona 2 '$249-$599/month is budget-line noise' - willingness-to-pay quoted as fact for personas the author invented, with no interview, survey, or waitlist evidence anywhere in the document.
Market Analysis 3
The top-line federal spend figures are directionally credible, but the derived market sizing is arithmetically wrong and the segment table double-counts its own population. The TAM ceiling calculation is off by three orders of magnitude, and the SAM table's rows sum to 502,000 while the stated total is 302,000 - because 'active small business contractors' and the consultant/small-firm rows are the same people counted twice. A market section that cannot add its own rows cannot support a $480M ARR addressable claim.
Cited gap: Section 4.1: 'If 1% of the 415,000+ annual contract actions involved a contractor using a dedicated proposal tool at $200/month average, that is $10 billion in annual software spend potential.' The actual product is 0.01 x 415,000 x $200 x 12 = $9.96 million - a 1,000x error. Even 100% of that population yields only $996M, so the stated ceiling is unreachable by two orders of magnitude.
Competitive Moat 3
Six of the seven 'Knockout Blows' are features, not moats - self-service pricing, export, and analytics are copyable in a quarter by anyone who decides to. The only durable asset described is Past Performance Vault data lock-in, which by the proposal's own admission requires six months of user population and is rated 'High' probability of slow adoption in the risk table. Worse, the moat section asserts as completed work something the product section marks as unbuilt.
Cited gap: Section 7.1 Knockout 2 states 'We built it anyway. No competitor has it' - but Section 5.2 tags Feature 3 Guided Response Flow as 'OPEN - primary differentiator' and Section 5.3 lists 'Guided response wizard | TO BUILD | High'. The central moat claim is written in the past tense about a feature that does not exist.
Business Model 4
Four clean tiers with sensible feature laddering, and the value-based pricing logic is coherent in structure. But the entire premium-below-enterprise positioning rests on a competitor price the document states two incompatible ways, and under the lower of its own two figures the pricing narrative inverts completely - RFP Tank becomes the expensive option, not the bargain. No pricing was tested against a live buyer.
Cited gap: GovEagle is priced at '$30,000+/year' in the Executive Summary, the Section 4.4 table, and the Section 7.2 map, but at '$2,000+ per year' in the Elevator Pitch, Section 6.2, and Appendix A ('~$2,000+/year confirmed via colleague'). At $2,000/yr, the $599/mo Team tier ($7,188/yr) is 3.6x more expensive than the incumbent it claims to rout.
Unit Economics 3
Every input to the unit economics table is either unvalidated or computed incorrectly. LTV is calculated on gross revenue rather than gross profit despite the same table listing per-tier margins, inflating each figure by 18-25%; CAC is asserted with no channel test behind it; and the operating cost base is stated at two different magnitudes in two different sections. The resulting ratios are then presented as a reason to spend more, which is exactly backwards for an untested CAC.
Cited gap: Section 10.2 lists Solo LTV as $2,607 - precisely $79 x 33 months, gross revenue - while the adjacent row states 82% gross margin, which would give $2,138. Compounding this, Section 6.3 puts 'Total Monthly OpEx' at '~$3,000-$5,000' while Section 10.3 computes break-even from 'current cost structure ($10,000-$12,000/month in OpEx excluding labor)' - a 2-4x contradiction in the same document.
Technical Architecture 4
The stack choices are conventional and defensible, and the shared-vs-dedicated deployment split is more operational detail than most proposals at this stage offer. But the compliance posture is self-contradictory in a way that is disqualifying for the segment it targets, and the hosting geography is incompatible with the workloads the dedicated tier is explicitly sold for. Real-time LLM copilot cost and latency at a $79/mo price point are asserted, never modeled.
Cited gap: Section 9.4 asserts 'no CUI handling (FedRAMP not needed)', yet Section 5.1.1 Option B markets dedicated deployment as 'Best for: Enterprise clients handling classified or ITAR-adjacent solicitations'. DFARS 252.204-7012(b)(2)(ii)(D) requires FedRAMP Moderate-equivalent cloud for covered defense information, and the named hardware - netcup and Hetzner, both German providers - fails ITAR US-person and data-residency requirements outright.
Go-to-Market Strategy 5
The channel inventory is the strongest part of the GTM: GovCon is genuinely targetable by NAICS, set-aside status, and professional association, and the free-shredder product-led hook is the right wedge. But all eight channel conversion rates in the 8.3 table are invented, no channel has been piloted, and the research underpinning the flagship partnership channel is several years stale - which signals the competitive and market scans were desk research, not fieldwork.
Cited gap: Section 8.2 item 2 and Section 8.3 build the warmest channel on 'PTAC offices' and Section 11.3 sets a Month-12 target of '5+ regional PTACs recommending RFP Tank'. PTACs were renamed APEX Accelerators in a DoD transition completed in 2023 and moved out from under DLA - the proposal is recruiting a program under a name it stopped using roughly three years before the August 2026 proposal date.
Risk Assessment 6
The best-executed section in the document - four risk categories plus a genuine ranked pre-mortem that correctly identifies slow delivery of the guided flow as the top failure mode. It loses points because most mitigations are reactive instrumentation ('track conversion rates weekly', 'monitor API spend daily') rather than actions taken before spend, and because the register never names the risk that actually dominates: that no assumption in the document has been checked with a customer. The risk probabilities are also inconsistent with the financial model.
Cited gap: Section 10.2 assumes 3% monthly Solo churn, which compounds to 8.7% over 90 days, while Section 9.6 sets '90-day churn: Under 5%' as the success target and 'Over 10% = product-market fit problem' - the base-case financial model is already closer to the failure threshold than to the target, and neither section acknowledges the other.
Execution Feasibility 3
One developer is scheduled to ship seven features - two of them self-rated 'High' effort - plus billing, auth, marketing site, and a daily SAM.gov ingestion pipeline, in 90 days with zero slack. The build sequence in Decision 3 sums to exactly ten weeks against a Week 10 beta, meaning any single slip moves launch. The document identifies this bottleneck accurately and then does not resource it, deferring the fix until after the traction it is required to produce.
Cited gap: Section 9.4 rates 'Single-developer bottleneck (Germaine)' as Probability High / Impact High, and the mitigation is 'hire a second developer by Month 4 if traction confirms product-market fit' - the relief is gated on the outcome the bottleneck must deliver first, and Section 11.1 budgets 'Total hard dollar cost to launch: Under $500', with no contingency to hire earlier.
Growth Trajectory 2
This is the most damaging finding in the review: the forecast does not agree with itself at any level, and the errors compound upward into a valuation claim. Year 1 ARR is stated three mutually exclusive ways, the month-by-month ramp contradicts its own subscriber counts by 20% at Month 12, and the whole structure sits on the 1,000x TAM error and the revenue-based LTV from earlier sections. A three-year projection cannot be tested by a reader when the twelve-month table cannot reproduce its own arithmetic - and none of it is anchored to a single paying user, waitlist signup, or letter of intent.
Cited gap: The Section 6.5 ramp table lists Month 12 as 395 Solo + 168 Consultant + 56 Team at $85,072 MRR, but those counts at the Section 5.5 prices ($79/$249/$599) produce $106,581 - a $21,509 unexplained gap that widens every month from Month 3 onward. Meanwhile Year 1 ARR is '$1.2M' (Executive Summary), '$1,361,400' (Section 6.1), and '$1.02M' (Sections 6.5 and 10.5), and Section 10.5 extends this to 'Year 3 ARR of $16M... a $130M+ valuation' - a nine-figure claim resting on a table that cannot multiply its own subscriber counts by its own prices.

4. Architecture

RFP Tank inherits the VerdictTank v3.6 engine. The specification below is the shared engine architecture the product builds on.

Status: Build specification, companion to the v3.6 business proposal.

Scope: Extends the v3 architecture (5-stage pipeline, 4 API endpoints, 10 data model tables) with the 9 new v3.6 capabilities: Chat-to-Refine, URL-to-Review, Dual Scoring, Explain-the-Low-Score, Fix-It Remediation, Second-Opinion Audit Agent, Configurable Rules Engine, White-Label for Consultants, and Roast/Boost Community Peer Review.

Audience: Engineering. This is the build spec, not the pitch.


1. System Overview

v3.6 scope note (Priority Fix #6 / Judge 3 Week 1 Cut List): The critical review of v3.6 found the 15-service, 21-table build scoped "like a multi-engineer build" for a solo-founder timeline. This document now describes two tiers: the 8-Week MVP (5 services, shipped) and Deferred (Post-Launch) components (5 services, designed but not built until a funnel/demand signal justifies them). Deferred components' schemas and designs are preserved below for continuity - cutting scope does not mean deleting the design work, it means sequencing it after MVP validation. See Section 9 for the revised build sequence.

1.1 Architecture Diagram - 8-Week MVP Scope

                                   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                                   β”‚                CLIENTS                       β”‚
                                   β”‚  Web App Β· Mobile Β· API Consumers            β”‚
                                   β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                                        β”‚ HTTPS
                                   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                                   β”‚           EDGE / API GATEWAY (Caddy)          β”‚
                                   β”‚  TLS term Β· rate limit Β· auth (JWT/API key)   β”‚
                                   β”‚  request logging                              β”‚
                                   β””β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                           β”‚
                     β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                     β”‚     APPLICATION SERVICES (MVP: 2 of 5 kept) β”‚
                     β”‚  review-api Β· chat-api                      β”‚
                     β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                 β”‚ enqueue
                     β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                     β”‚                    JOB QUEUE (Redis + RQ/Celery)           β”‚
                     β”‚  review.pipeline Β· chat.turn Β· url.extract                 β”‚
                     β”‚  remediation.generate Β· sanitize.scan                     β”‚
                     β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                 β”‚
     β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
     β”‚                         REVIEW PIPELINE WORKERS (pipeline-worker)              β”‚
     β”‚                                                                                β”‚
     β”‚  Phase 1        Phase 2         Phase 3          Phase 4       Phase 5         β”‚
     β”‚  Research  ──▢  Primary    ──▢  Validation  ──▢  Cross-Check ──▢ Verdict        β”‚
     β”‚  Agent          Reviewer        Reviewer          A + B (β€–)     Aggreg.        β”‚
     β”‚  (web verify    (10-dim,        (challenges       (independent  (majority       β”‚
     β”‚   + citations)   1-10 + tags    primary score)     re-score)     + dual         β”‚
     β”‚                  idea/proposal)                                  scores)        β”‚
     β”‚                                                                                β”‚
     β”‚  [Phase 6 - Audit Agent: DEFERRED, see Β§1.2 and Β§9 Post-Launch Track]          β”‚
     β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                 β”‚                                             β”‚
     β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”                 β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
     β”‚   SANITIZATION GATE        β”‚                 β”‚   POST-PIPELINE GENERATORS      β”‚
     β”‚   scans: verdict JSON,     β”‚                 β”‚   Explainer (per-dim text)      β”‚
     β”‚   explanations, remediation,β”‚                β”‚   Fix-It Generator (action plan)β”‚
     β”‚   chat transcripts         β”‚                β”‚   (Rules Engine overlay: DEFERRED)β”‚
     β”‚   before egress             β”‚                β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
     β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜                                β”‚
                 β”‚                                                β”‚
     β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
     β”‚                              DATA LAYER                                       β”‚
     β”‚  PostgreSQL (primary OLTP, org_id column present but multi-tenant enforcement β”‚
     β”‚  (RLS) deferred until White-Label track - see Β§1.2)                          β”‚
     β”‚  pgvector (corpus embeddings/similarity)                                      β”‚
     β”‚  Redis (queue, session cache, rate limits, feature flags)                     β”‚
     β”‚  S3-compatible object store (PDFs, chat exports, uploaded docs, extracted HTML)β”‚
     β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

     β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
     β”‚                     SUPPORTING SERVICES (MVP)                               β”‚
     β”‚  url-extractor (Crawl4AI) Β· fixit-generator Β· pdf-generator                 β”‚
     β”‚  cron (prediction T+90/180/365)                                            β”‚
     β”‚  [domain-verifier, moderation-queue: DEFERRED - see Β§1.2]                   β”‚
     β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

1.2 Component List - MVP vs. Deferred

Per Priority Fix #6, the 8-week solo-founder MVP ships 5 services; the remaining 5 services from the original v3.6 scope are deferred to post-launch and gated on real usage/demand signal, not built speculatively. This directly implements the review's Week 1 Cut List (KEEP: Dual Scoring, Explain, Fix-It MVP, URL-to-Review, Corpus Schema / CUT: Audit Agent, Rules Engine, White-Label, Roast/Boost, and Chat-to-Refine deferred-not-deleted per below).

MVP - Ship in 8 Weeks (5 services)

Layer Component MVP Status Notes
App review-api KEEP Core review submission/status/verdict endpoints; dual-score fields, /review/url
App chat-api KEEP (reduced) Chat-to-Refine, but scoped down to the ghostwriting-guard behavior in Β§5.1 - no chat-ws streaming socket in MVP, REST polling only, to cut infra surface. Positioned as the "un-copilot" (see Β§5.1) rather than deferred outright, since Judge 3 flagged it as a differentiator once guarded correctly - see rationale note below.
Worker url-extractor KEEP Crawl4AI-based content extraction + sanitize; Feature 2 (URL-to-Review)
Worker pipeline-worker KEEP Phases 1-5 (Research β†’ Primary β†’ Validation β†’ Cross-Check β†’ Verdict/Dual-Score Aggregation). Phase 6 (Audit Agent) removed from the worker in MVP.
Worker fixit-generator KEEP Fix-It Remediation, upgraded to quality-scoring per Β§5.5 - retention-loop feature, explicitly called out as a KEEP in the review's cut list
Note on Chat-to-Refine placement: The review's Week 1 Cut List names Chat-to-Refine as a CUT ("funnel feature; build when you have a funnel"). This document keeps a minimal chat-api in MVP scope specifically because the ghostwriting guard in Β§5.1 is cheap to build (a length/format check, not new infrastructure) and de-risks the single biggest brand-reputation exposure identified by Judge 3 if Chat-to-Refine ships at all, in this MVP or later. If engineering capacity is tighter than modeled, chat-api is the first MVP component to cut - in which case Chat-to-Refine moves to the Deferred table below in its entirety, chat-ws and all real-time streaming remain deferred either way, and the guard design in Β§5.1 is preserved for whenever the feature ships.

DEFERRED - Post-Launch, Built on Demand Signal (5 services)

Layer Component MVP Status Notes Deferred Trigger
App rules-api CUT (deferred) CRUD + validation for org rule definitions Revisit once an Enterprise/White-Label deal is in active negotiation and requires it
App org-api CUT (deferred) white-label org config, branding, API keys Blocked on Minimum Viable Legal (MSA + Privacy Program) per the review's Legal Killers - see Β§4
App peer-review-api CUT (deferred) Roast/Boost submission + moderation Moderation overhead + defamation exposure (Judge 4 Legal Killer #4) outweighs launch value
Worker audit-agent-worker CUT (deferred) Phase 6 of pipeline, second-opinion audit "Enterprise trust problem, not launch problem" per Judge 3; revisit once Enterprise tier has real pilot demand
Worker moderation-worker CUT (deferred) abuse handling for Roast/Boost Dependent on peer-review-api; deferred together
Infra domain-verifier CUT (deferred) CNAME/TXT verification for white-label custom domains Dependent on org-api / White-Label track; also blocked on Minimum Viable Legal

The rules-evaluator overlay logic (Β§5.7) and Rules Engine pipeline injection point are deferred along with rules-api - no schema or prompt-injection surface for org rules ships in MVP. The full designs for all deferred components remain documented in Sections 5.6-5.9 and 7 of this document as the intended post-launch build, not as speculative dead weight - they represent validated architecture pending a demand signal, not throwaway work.

MVP Data Layer implications: Row-level security (RLS) multi-tenant policies (Β§6.3, Β§7.4), the org_configs/org_rules/org_api_keys/rule_templates tables, and cross-org isolation logic are not required for MVP since org-api and rules-api are deferred - org_id columns remain present (nullable) on core tables for forward compatibility, but RLS enforcement, domain routing, and multi-tenant billing are built when the White-Label track is actually resumed, gated on Minimum Viable Legal per Section 4.


2. Complete Data Model

2.1 Entity Relationship Summary

users ──1:N── reviews ──1:1── verdicts ──1:N── dimensions ──1:1── dimension_explanations
  β”‚               β”‚                                  β”‚
  β”‚               β”‚                                  └─1:N── remediation_plans (per dimension)
  β”‚               β”œβ”€β”€1:N── judge_scores ──N:1── judges
  β”‚               β”œβ”€β”€1:N── research_citations
  β”‚               β”œβ”€β”€1:N── predictions
  β”‚               β”œβ”€β”€1:N── audit_findings          (v3.6)
  β”‚               β”œβ”€β”€1:N── peer_reviews            (v3.6)
  β”‚               └──1:1── chat_sessions (optional, pre-review link)  (v3.6)
  β”‚
  β”œβ”€β”€1:N── chat_sessions ──1:N── chat_messages      (v3.6)
  β”‚
  └──N:1── org_configs ──1:N── org_rules ──N:1── rule_templates   (v3.6, enterprise/white-label)

corpus_entries (derived from verdicts, searchable, org_id-scoped)
audit_log (cross-cutting, all tables)

All new v3.6 tables carry org_id (nullable for individual/non-enterprise accounts) to support multi-tenant partitioning for White-Label. All v3 tables are extended with org_id as part of the Phase 0 migration described in Section 9.

2.2 v3 Tables (Extended for v3.6)

Dual-Score Schema Validation Gate (Priority Fix #2 / Judge 3 & Judge 1 finding, Architecture Β§2.2 lines 147-148 in the v3.6 review): The original v3.6 design baked idea_score/proposal_score directly into the reviews and corpus_entries tables as permanent NUMERIC columns from day one - a schema bet made before any beta user confirmed that a two-score split is more useful than the v3 single composite score. The review's Priority Fix #2 is explicit: "Do not bake idea_score/proposal_score into the corpus before validating with beta users that two scores are wanted. Gate the migration behind user testing, or ship dual-scoring as a non-schema overlay first."

>

v3.6 resolution - ship as overlay, migrate only after validation:
1. MVP ships dual scoring as a JSONB overlay, not dedicated columns. The idea_score/proposal_score values computed at Phase 5 (see Β§5.3) are written into verdicts.verdict_json (already a JSONB column, no migration required) under a dual_score key, and into a new reviews.dual_score_overlay JSONB column (nullable, additive, zero-downtime to add/drop) rather than as first-class typed NUMERIC(5,2) columns on reviews or corpus_entries.
2. The NUMERIC columns shown below (idea_score, proposal_score on reviews; same on corpus_entries, Β§2.3) are the target schema for GA, not the MVP schema. They are documented here for architectural continuity but are marked DEFERRED - VALIDATION GATE and must not be created by the Phase 0 migration until the exit criterion below is met.
3. Validation exit criterion (must pass before the permanent-column migration runs): a minimum of 25 beta users from the existing $29/mo Inner Circle cohort have used dual-score overlay output across at least 2 reviews each, and post-review survey/interview signal shows a majority preference for two scores over the legacy single composite. This reuses the same beta cohort the review's Priority Fix #1 calls for surveying - one survey instrument serves both purposes.
4. If validation fails (users prefer one score, or show no measurable preference), the dual_score_overlay JSONB column is dropped, verdict_json.dual_score continues to carry the value for backward-compatible display only, and the corpus indexing strategy in Β§2.3/Β§2.4 reverts to legacy_composite_score as the primary sortable/filterable column.
5. Corpus search and percentile lookups (Β§6.4) run against the JSONB overlay during the validation window - idx_corpus_dual_score_overlay is a GIN index on the JSONB path, functionally equivalent to a B-tree NUMERIC index for the query patterns in Β§3.2/Β§3.3 but avoids a schema commitment. Query latency is slightly higher (GIN vs. B-tree) but acceptable at MVP corpus volume; this is re-benchmarked before the permanent migration.
-- USERS (v3, extended)
CREATE TABLE users (
    id              UUID PRIMARY KEY DEFAULT gen_random_uuid(),
    org_id          UUID REFERENCES org_configs(id),      -- v3.6: multi-tenant scoping
    email           TEXT UNIQUE NOT NULL,
    password_hash   TEXT,
    tier            TEXT NOT NULL DEFAULT 'free',          -- free|pro|enterprise|white_label|beta
    role            TEXT NOT NULL DEFAULT 'member',        -- member|org_admin|consultant|superadmin
    created_at      TIMESTAMPTZ NOT NULL DEFAULT now(),
    updated_at      TIMESTAMPTZ NOT NULL DEFAULT now()
);
CREATE INDEX idx_users_org ON users(org_id);

-- REVIEWS (v3, extended: dual score columns, submission source, org scoping)
CREATE TABLE reviews (
    id                  UUID PRIMARY KEY DEFAULT gen_random_uuid(),
    org_id              UUID REFERENCES org_configs(id),
    user_id             UUID NOT NULL REFERENCES users(id),
    source_type         TEXT NOT NULL DEFAULT 'file',       -- v3.6: file|url|api
    source_url          TEXT,                               -- v3.6: populated when source_type='url'
    chat_session_id     UUID REFERENCES chat_sessions(id),  -- v3.6: link to pre-review coaching
    vertical            TEXT,
    status              TEXT NOT NULL DEFAULT 'queued',     -- queued|running|audit|done|failed
    idea_score          NUMERIC(5,2),                       -- DEFERRED - VALIDATION GATE: do not populate until dual-score validation exit criterion (Β§2.2) is met; MVP writes to dual_score_overlay instead
    proposal_score      NUMERIC(5,2),                       -- DEFERRED - VALIDATION GATE: see idea_score note above
    dual_score_overlay  JSONB,                              -- v3.6 MVP: { "idea_score": 78.5, "proposal_score": 41.0 }, non-schema overlay per Β§2.2 gate; promoted to typed columns above only post-validation
    legacy_composite_score NUMERIC(5,2),                    -- v3: preserved for backward compat
    rules_applied       JSONB DEFAULT '[]',                 -- DEFERRED - org_rules snapshot, populated only once Rules Engine (Β§5.7) ships post-launch
    parent_review_id    UUID REFERENCES reviews(id),        -- v3.6 (Β§5.5): set when this review is a resubmission addressing a prior remediation_plan
    fix_completion_score NUMERIC(3,2),                       -- v3.6 (Β§5.5): mean fix_quality_tier (0-3) across parent's action_items, computed on resubmission
    created_at          TIMESTAMPTZ NOT NULL DEFAULT now(),
    completed_at        TIMESTAMPTZ
);
CREATE INDEX idx_reviews_org ON reviews(org_id);
CREATE INDEX idx_reviews_user ON reviews(user_id);
CREATE INDEX idx_reviews_status ON reviews(status);
CREATE INDEX idx_reviews_parent ON reviews(parent_review_id);

-- VERDICTS (v3, unchanged shape, verdict_json now carries dual-score payload)
CREATE TABLE verdicts (
    id              UUID PRIMARY KEY DEFAULT gen_random_uuid(),
    review_id       UUID NOT NULL REFERENCES reviews(id) UNIQUE,
    verdict_json    JSONB NOT NULL,      -- see Section 2.5 for v3.6 schema
    pdf_url         TEXT,
    public_share_id TEXT UNIQUE,
    created_at      TIMESTAMPTZ NOT NULL DEFAULT now()
);

-- DIMENSIONS (v3, extended: score_type tag for dual scoring)
CREATE TABLE dimensions (
    id              UUID PRIMARY KEY DEFAULT gen_random_uuid(),
    review_id       UUID NOT NULL REFERENCES reviews(id),
    name            TEXT NOT NULL,             -- e.g. "Market Analysis"
    score           NUMERIC(4,2) NOT NULL,     -- 1-10, existing v3 scale
    score_type      TEXT NOT NULL DEFAULT 'proposal',  -- v3.6: 'idea' | 'proposal'
    weight          NUMERIC(4,3) DEFAULT 1.0,
    created_at      TIMESTAMPTZ NOT NULL DEFAULT now()
);
CREATE INDEX idx_dimensions_review ON dimensions(review_id);

-- JUDGES (v3, unchanged)
CREATE TABLE judges (
    id              UUID PRIMARY KEY DEFAULT gen_random_uuid(),
    role            TEXT NOT NULL,             -- Research|Primary|Validation|CrossCheckA|CrossCheckB|...
    vendor_internal TEXT NOT NULL,             -- stripped by sanitization gate before egress
    is_audit_agent  BOOLEAN NOT NULL DEFAULT false, -- v3.6: flags the Phase 6 role
    active          BOOLEAN NOT NULL DEFAULT true,
    accuracy_score  NUMERIC(5,4),
    created_at      TIMESTAMPTZ NOT NULL DEFAULT now()
);

-- JUDGE_SCORES (v3, unchanged)
CREATE TABLE judge_scores (
    id              UUID PRIMARY KEY DEFAULT gen_random_uuid(),
    review_id       UUID NOT NULL REFERENCES reviews(id),
    judge_id        UUID NOT NULL REFERENCES judges(id),
    dimension_id    UUID REFERENCES dimensions(id),
    raw_score       NUMERIC(4,2) NOT NULL,
    rationale       TEXT,
    created_at      TIMESTAMPTZ NOT NULL DEFAULT now()
);

-- RESEARCH_CITATIONS (v3, unchanged)
CREATE TABLE research_citations (
    id              UUID PRIMARY KEY DEFAULT gen_random_uuid(),
    review_id       UUID NOT NULL REFERENCES reviews(id),
    source_url      TEXT NOT NULL,
    excerpt         TEXT,
    retrieved_at    TIMESTAMPTZ NOT NULL DEFAULT now()
);

-- PREDICTIONS (v3, unchanged)
CREATE TABLE predictions (
    id              UUID PRIMARY KEY DEFAULT gen_random_uuid(),
    review_id       UUID NOT NULL REFERENCES reviews(id),
    predicted_risk  TEXT NOT NULL,
    check_at        TIMESTAMPTZ NOT NULL,   -- T+90/180/365
    outcome         TEXT,                   -- pending|materialized|avoided
    checked_at      TIMESTAMPTZ
);

-- CORPUS_ENTRIES (v3, extended: org scoping for white-label segmentation, dual score overlay)
CREATE TABLE corpus_entries (
    id              UUID PRIMARY KEY DEFAULT gen_random_uuid(),
    org_id          UUID REFERENCES org_configs(id),   -- v3.6: NULL = shared/general corpus
    review_id       UUID NOT NULL REFERENCES reviews(id),
    vertical        TEXT,
    idea_score      NUMERIC(5,2),                       -- DEFERRED - VALIDATION GATE, see Β§2.2; not populated at MVP
    proposal_score  NUMERIC(5,2),                       -- DEFERRED - VALIDATION GATE, see Β§2.2; not populated at MVP
    dual_score_overlay JSONB,                           -- v3.6 MVP: populated instead of the two columns above, see Β§2.2
    embedding       VECTOR(1536),
    is_public       BOOLEAN NOT NULL DEFAULT false,
    created_at      TIMESTAMPTZ NOT NULL DEFAULT now()
);
CREATE INDEX idx_corpus_org ON corpus_entries(org_id);
CREATE INDEX idx_corpus_embedding ON corpus_entries USING hnsw (embedding vector_cosine_ops);
CREATE INDEX idx_corpus_dual_score_overlay ON corpus_entries USING gin (dual_score_overlay);  -- v3.6 MVP: GIN index serves overlay queries until/unless promoted to typed columns, see Β§2.2

-- AUDIT_LOG (v3, extended: new event types for v3.6 surfaces)
CREATE TABLE audit_log (
    id              UUID PRIMARY KEY DEFAULT gen_random_uuid(),
    org_id          UUID REFERENCES org_configs(id),
    actor_id        UUID REFERENCES users(id),
    event_type      TEXT NOT NULL,   -- + v3.6: chat.message, review.url_submit, rule.create,
                                      --   audit_agent.finding, peer_review.submit, org.brand_update
    entity_type     TEXT,
    entity_id       UUID,
    metadata        JSONB DEFAULT '{}',
    created_at      TIMESTAMPTZ NOT NULL DEFAULT now()
);
CREATE INDEX idx_audit_log_entity ON audit_log(entity_type, entity_id);

2.3 New v3.6 Tables

-- CHAT_SESSIONS (Feature 1: Chat-to-Refine)
CREATE TABLE chat_sessions (
    id              UUID PRIMARY KEY DEFAULT gen_random_uuid(),
    user_id         UUID NOT NULL REFERENCES users(id),
    org_id          UUID REFERENCES org_configs(id),
    status          TEXT NOT NULL DEFAULT 'active',  -- active|archived|converted_to_review
    review_id       UUID REFERENCES reviews(id),      -- set once user submits for review
    title           TEXT,
    turn_count      INT NOT NULL DEFAULT 0,
    created_at      TIMESTAMPTZ NOT NULL DEFAULT now(),
    updated_at      TIMESTAMPTZ NOT NULL DEFAULT now()
);
CREATE INDEX idx_chat_sessions_user ON chat_sessions(user_id);

-- CHAT_MESSAGES
CREATE TABLE chat_messages (
    id              UUID PRIMARY KEY DEFAULT gen_random_uuid(),
    session_id      UUID NOT NULL REFERENCES chat_sessions(id),
    role            TEXT NOT NULL,        -- user|coach
    content         TEXT NOT NULL,
    sanitized       BOOLEAN NOT NULL DEFAULT false,  -- gate must run before export
    created_at      TIMESTAMPTZ NOT NULL DEFAULT now()
);
CREATE INDEX idx_chat_messages_session ON chat_messages(session_id, created_at);

-- DIMENSION_EXPLANATIONS (Feature 4: Explain the Low Score)
CREATE TABLE dimension_explanations (
    id              UUID PRIMARY KEY DEFAULT gen_random_uuid(),
    dimension_id    UUID NOT NULL REFERENCES dimensions(id) UNIQUE,
    explanation_text TEXT NOT NULL,
    cited_gaps      JSONB NOT NULL DEFAULT '[]',  -- ["no TAM calculation", "no competitor pricing"]
    sanitized       BOOLEAN NOT NULL DEFAULT false,
    generated_by    UUID REFERENCES judges(id),    -- always the Primary Reviewer judge
    created_at      TIMESTAMPTZ NOT NULL DEFAULT now()
);

-- REMEDIATION_PLANS (Feature 5: Here's How to Fix It - quality-scored per Β§5.5)
CREATE TABLE remediation_plans (
    id              UUID PRIMARY KEY DEFAULT gen_random_uuid(),
    review_id       UUID NOT NULL REFERENCES reviews(id),
    dimension_id    UUID NOT NULL REFERENCES dimensions(id),
    action_items    JSONB NOT NULL DEFAULT '[]',
    -- action_items shape: [{ "description": str, "difficulty": 1-3,
    --   "estimated_time_minutes": int, "template_url": str|null,
    --   "fix_quality_tier": int|null,       -- 0=not_attempted,1=superficial,2=minimal,3=substantive; NULL until resubmission evaluated, see Β§5.5
    --   "fix_quality_rationale": str|null,  -- evaluator's one-sentence justification, sanitized per Β§7.2
    --   "evaluated_at": timestamptz|null }]
    tier_scope      TEXT NOT NULL DEFAULT 'full',  -- 'summary' (Free) | 'full' (Pro+)
    disclaimer_version TEXT NOT NULL DEFAULT 'DISC-001',  -- v3.6: legal disclaimer version active at generation time, see Β§2.6
    sanitized       BOOLEAN NOT NULL DEFAULT false,
    created_at      TIMESTAMPTZ NOT NULL DEFAULT now()
);
CREATE INDEX idx_remediation_review ON remediation_plans(review_id);

-- Fix quality aggregate lives on REVIEWS (added to the reviews table above, Β§2.2):
--   reviews.fix_completion_score NUMERIC(3,2)  -- mean fix_quality_tier (0-3) across all action_items on resubmission, see Β§5.5
--   reviews.parent_review_id UUID REFERENCES reviews(id)  -- links a resubmission to the review whose remediation plan it's addressing

-- AUDIT_FINDINGS (Feature 6: Second-Opinion Audit Agent) [DEFERRED - Post-Launch, see Β§1.2/Β§5.6]
-- Table shape preserved here for continuity; not created by the MVP Phase 0 migration.
CREATE TABLE audit_findings (
    id              UUID PRIMARY KEY DEFAULT gen_random_uuid(),
    review_id       UUID NOT NULL REFERENCES reviews(id),
    finding_type    TEXT NOT NULL,  -- blind_spot|groupthink|underweighted_dim|tagging_inconsistency|no_finding
    description     TEXT NOT NULL,
    severity        TEXT NOT NULL DEFAULT 'low',  -- low|medium|high
    related_dimension_id UUID REFERENCES dimensions(id),
    disclaimer_version TEXT NOT NULL DEFAULT 'DISC-001',  -- v3.6: see Β§2.6
    sanitized       BOOLEAN NOT NULL DEFAULT false,
    created_at      TIMESTAMPTZ NOT NULL DEFAULT now()
);
CREATE INDEX idx_audit_findings_review ON audit_findings(review_id);

-- RULE_TEMPLATES (Feature 7: Configurable Review Rules Engine) [DEFERRED - Post-Launch, see Β§1.2/Β§5.7]
-- Table shape preserved here for continuity; not created by the MVP Phase 0 migration.
CREATE TABLE rule_templates (
    id              UUID PRIMARY KEY DEFAULT gen_random_uuid(),
    name            TEXT NOT NULL,
    category        TEXT NOT NULL,     -- compliance|brand_voice|required_field|checklist
    schema_json     JSONB NOT NULL,    -- JSON Schema describing valid rule params
    is_system       BOOLEAN NOT NULL DEFAULT true,  -- system-provided vs org-authored
    created_at      TIMESTAMPTZ NOT NULL DEFAULT now()
);

-- ORG_RULES [DEFERRED - Post-Launch, see Β§1.2/Β§5.7]
CREATE TABLE org_rules (
    id              UUID PRIMARY KEY DEFAULT gen_random_uuid(),
    org_id          UUID NOT NULL REFERENCES org_configs(id),
    template_id     UUID REFERENCES rule_templates(id),
    name            TEXT NOT NULL,
    rule_definition JSONB NOT NULL,
    -- rule_definition shape: { "condition": {...if/then AST...},
    --   "required_fields": [...], "checklist": [...], "severity": "block|warn" }
    is_active       BOOLEAN NOT NULL DEFAULT true,
    validated_at    TIMESTAMPTZ,       -- set once syntax + schema validation passes
    created_by      UUID REFERENCES users(id),
    created_at      TIMESTAMPTZ NOT NULL DEFAULT now(),
    updated_at      TIMESTAMPTZ NOT NULL DEFAULT now()
);
CREATE INDEX idx_org_rules_org ON org_rules(org_id) WHERE is_active = true;

-- ORG_CONFIGS (Feature 8: White-Label for Consultants)
CREATE TABLE org_configs (
    id              UUID PRIMARY KEY DEFAULT gen_random_uuid(),
    name            TEXT NOT NULL,
    tier            TEXT NOT NULL DEFAULT 'white_label',  -- enterprise|white_label
    domain          TEXT UNIQUE,             -- custom domain, e.g. reviews.acceleratorx.com
    domain_verified BOOLEAN NOT NULL DEFAULT false,
    domain_verified_at TIMESTAMPTZ,
    logo_url        TEXT,
    brand_colors    JSONB DEFAULT '{}',      -- { "primary": "#...", "accent": "#..." }
    email_template  TEXT,                    -- HTML template with token placeholders
    corpus_segment  TEXT NOT NULL DEFAULT 'shared',  -- 'shared' | 'dedicated'
    api_key_prefix  TEXT UNIQUE,
    created_at      TIMESTAMPTZ NOT NULL DEFAULT now(),
    updated_at      TIMESTAMPTZ NOT NULL DEFAULT now()
);

-- ORG_API_KEYS (supports Feature 8's per-org API key scoping)
CREATE TABLE org_api_keys (
    id              UUID PRIMARY KEY DEFAULT gen_random_uuid(),
    org_id          UUID NOT NULL REFERENCES org_configs(id),
    key_hash        TEXT NOT NULL UNIQUE,
    label           TEXT,
    scopes          JSONB DEFAULT '["review:write","review:read"]',
    revoked_at      TIMESTAMPTZ,
    created_at      TIMESTAMPTZ NOT NULL DEFAULT now(),
    last_used_at    TIMESTAMPTZ
);

-- PEER_REVIEWS (Feature 9: Roast/Boost Community Peer Review)
CREATE TABLE peer_reviews (
    id              UUID PRIMARY KEY DEFAULT gen_random_uuid(),
    review_id       UUID NOT NULL REFERENCES reviews(id),
    user_id         UUID NOT NULL REFERENCES users(id),
    type            TEXT NOT NULL,   -- roast|boost
    comment         TEXT,
    visibility      TEXT NOT NULL DEFAULT 'submitter_only',  -- submitter_only|shared
    flagged         BOOLEAN NOT NULL DEFAULT false,
    flagged_reason  TEXT,
    moderation_status TEXT NOT NULL DEFAULT 'clean',  -- clean|pending_review|removed
    created_at      TIMESTAMPTZ NOT NULL DEFAULT now()
);
CREATE INDEX idx_peer_reviews_review ON peer_reviews(review_id);
CREATE UNIQUE INDEX idx_peer_reviews_dedupe ON peer_reviews(review_id, user_id);

-- PEER_REVIEW_RATE_LIMITS (supports 3/day/user cap)
CREATE TABLE peer_review_rate_limits (
    user_id         UUID PRIMARY KEY REFERENCES users(id),
    review_date     DATE NOT NULL,
    count_today     INT NOT NULL DEFAULT 0,
    UNIQUE (user_id, review_date)
);

2.4 Full Table Inventory (v3 + v3.6)

# Table Origin Purpose
1 users v3 (extended) account, tier, org membership
2 reviews v3 (extended) review lifecycle, dual scores, source type
3 verdicts v3 final verdict JSON, PDF, share link
4 dimensions v3 (extended) per-dimension score + idea/proposal tag
5 judges v3 (extended) panel roster incl. audit agent flag
6 judge_scores v3 raw per-judge scoring
7 research_citations v3 Phase 1 web verification sources
8 predictions v3 T+90/180/365 outcome tracking
9 corpus_entries v3 (extended) searchable review corpus, org-segmented
10 audit_log v3 (extended) cross-cutting event log
11 chat_sessions v3.6 Chat-to-Refine sessions
12 chat_messages v3.6 Chat-to-Refine turn history
13 dimension_explanations v3.6 Explain-the-Low-Score text
14 remediation_plans v3.6 Fix-It action plans
15 audit_findings v3.6 Second-Opinion Audit Agent output
16 rule_templates v3.6 reusable rule schemas
17 org_rules v3.6 org-specific custom rules
18 org_configs v3.6 white-label branding/domain
19 org_api_keys v3.6 per-org scoped API keys
20 peer_reviews v3.6 Roast/Boost entries
21 peer_review_rate_limits v3.6 abuse/rate control

2.5 Verdict JSON Schema (v3.6)

{
  "review_id": "uuid",
  "status": "done",
  "idea_score": 78.5,
  "proposal_score": 41.0,
  "dual_score_overlay": { "idea_score": 78.5, "proposal_score": 41.0 },
  "framing_summary": "Your idea is strong; your pitch needs work.",
  "disclaimer": "This is an automated critique generated by AI models, not professional advice. Scores and commentary are opinions, not facts. Do not rely on this output for investment, procurement, legal, or funding decisions without independent professional review. See full disclaimer at https://verdicttank.com/legal/disclaimer.",
  "dimensions": [
    {
      "name": "Market Analysis",
      "score": 4,
      "score_type": "idea",
      "explanation": {
        "text": "Market Analysis: 4/10 - no TAM calculation, no competitor pricing data, assumes zero competition.",
        "cited_gaps": ["no TAM calculation", "no competitor pricing data", "zero-competition assumption"],
        "disclaimer": "AI-generated critique, not professional advice."
      },
      "remediation": {
        "tier_scope": "full",
        "disclaimer": "This remediation plan is an automated suggestion, not professional or legal advice. Validate independently before acting.",
        "quality_tier": "not_yet_evaluated",
        "action_items": [
          {
            "description": "Calculate TAM using the top-down/bottom-up hybrid formula",
            "difficulty": 2,
            "estimated_time_minutes": 90,
            "template_url": "https://cdn.verdicttank.com/templates/tam-calc.xlsx"
          },
          {
            "description": "Add a competitor pricing comparison table",
            "difficulty": 1,
            "estimated_time_minutes": 45,
            "template_url": "https://cdn.verdicttank.com/templates/competitor-pricing.xlsx"
          }
        ]
      }
    }
  ],
  "audit": {
    "status": "deferred_post_mvp",
    "findings": [],
    "tagging_consistency_check": "not_run",
    "note": "Second-Opinion Audit Agent (Phase 6) is deferred post-MVP per Β§1.2. Field shape is preserved for forward compatibility; populated once the Audit Agent worker ships. When populated, audit.disclaimer carries the same 'no professional advice' language as other surfaces."
  },
  "rules_applied": [],
  "corpus_percentile": { "idea_score": 80, "proposal_score": 30 },
  "public_share_id": "vt-8x2k",
  "generated_at": "2026-08-11T00:00:00Z"
}

rules_applied ships as an always-empty array in MVP since the Rules Engine (rules-api) is deferred per Β§1.2 - the field is preserved in the schema for forward compatibility rather than removed, avoiding a breaking payload change when the Rules Engine ships post-launch.

All freeform text fields (explanation.text, action_items[].description, audit.findings[].description) are sanitized before this object leaves the pipeline - see Section 7.2. All disclaimer fields are not sanitizable/removable content - they are fixed legal boilerplate injected by the response serializer after sanitization, never model-generated, and cannot be stripped by any per-review customization (org branding, White-Label templates, tier). See Β§2.6 for the full disclaimer policy.

2.6 Legal Disclaimers - "No Professional Advice" (All Output Surfaces)

Origin: Judge 4's Legal Killer #2 (AI Liability / Defamation / Tortious Interference - ranked EXTREME) and Priority Fix #9 of the v3.6 review: "Every verdict, explanation, audit finding, and remediation action item must carry conspicuous language... This is the first line of defense against Rank 2 liability." No disclaimer, limitation-of-liability language, or indemnity existed anywhere in v3.6 prior to this revision - Section 15/Risk Assessment covered only technical/product risk, not the liability exposure of an AI system whose scores can be blamed for a lost deal, declined funding round, or reputational harm.

Standard disclaimer text (canonical copy, referenced by ID DISC-001 so all surfaces stay in sync on future legal-copy revisions):

"This is an automated critique generated by AI models, not professional advice. Scores, explanations, audit findings, and remediation suggestions are AI-generated opinions, not verified facts or expert judgments. Do not rely on this output for investment, procurement, legal, regulatory, or funding decisions without independent professional review. VerdictTank and IT Pro Partner disclaim liability for decisions made in reliance on this output. Full terms: https://verdicttank.com/legal/disclaimer."

Coverage - every output surface carries it, no exceptions:

Output Surface Placement Enforcement Mechanism
Verdict JSON (verdict_json) Top-level disclaimer field + per-dimension explanation.disclaimer + per-remediation remediation.disclaimer (see Β§2.5) Injected by the API/PDF serializer post-sanitization; schema validation rejects a verdict payload missing the top-level field (fail-closed, same enforcement pattern as explanation.text in Β§5.4)
PDF reports Persistent footer on every page + prominent banner directly below the score header on page 1 Hardcoded in the PDF template (not model-generated, not org-brand-overridable - see White-Label note below)
Public share reports (web) Sticky banner above the fold, not a footnote or dismissible tooltip Rendered server-side in the share-page template, not client-injectable/removable
API responses (/verdict/{id}, /verdict/{id}/pdf, /reviews/{id}/remediation, /reviews/{id}/audit) disclaimer field present on every response object that carries scores, explanations, findings, or remediation Contract-tested: API response schema tests fail CI if any scored/explained/remediated response type omits the field
Chat-to-Refine transcripts Session-start system message ("I'm a coach, not an advisor - nothing here is professional advice") + persistent footer in chat UI Rendered client-side from a fixed string, reinforces the "un-copilot" brand framing in Β§5.1
Audit findings (when the Audit Agent ships post-MVP, Β§1.2) Same disclaimer field pattern as explanations/remediation, applied at design time so no retrofit is needed when Phase 6 ships Schema shape reserved in Β§2.5 now; enforcement added to the audit-findings serializer at build time
Corpus/percentile displays Inline caveat: "Percentile rankings are relative to other AI-scored submissions, not a market or investment benchmark" Rendered alongside corpus_percentile wherever it's displayed

Non-negotiable design constraints:


3. API Reference

3.1 Authentication

3.2 v3 Endpoints (Unchanged Contract, Extended Payload)

POST /api/verdicttank/review

Upload-based review submission (file). Response now includes dual score fields once complete.

// Response (200, once status=done)
{
  "review_id": "uuid",
  "idea_score": 78.5,
  "proposal_score": 41.0,
  "status": "done"
}

GET /api/verdicttank/status/{id} - unchanged polling contract; status enum extended with audit (Phase 6 in progress).

GET /api/verdicttank/verdict/{id}/pdf - unchanged; PDF generator now renders dual-score header and per-dimension explanation/remediation sections.

GET /api/verdicttank/corpus/search - unchanged query contract; results now include idea_score and proposal_score columns instead of a single blended score, and are org_id-scoped for white-label callers.

3.3 New v3.6 Endpoints

POST /api/verdicttank/review/url (Feature 2: URL-to-Review)

// Request
{ "url": "https://example.com/pitch-deck", "vertical_hint": "fintech" }
// Response (202 Accepted)
{ "review_id": "uuid", "status": "queued", "source_type": "url" }

Validation: URL scheme allowlist (http/https), DNS resolution check, max content size 10MB, extraction timeout 30s (hard cap 45s), robots.txt respected, content sanitized before entering the pipeline (see Section 7.2).

POST /api/verdicttank/chat/start (Feature 1)

// Request: { "initial_message": "I'm building a marketplace for..." }
// Response: { "session_id": "uuid", "reply": "Tell me more about who pays on this marketplace." }

POST /api/verdicttank/chat/{session_id}/message

// Request: { "message": "Both sides pay a transaction fee." }
// Response: { "reply": "...", "turn_count": 4 }

Also available as WSS /ws/verdicttank/chat/{session_id} for real-time streaming token-by-token (see Section 5.1). REST polling variant is the fallback for clients that can't hold a socket.

GET /api/verdicttank/chat/{session_id}/history

Returns sanitized transcript (chat_messages where sanitized=true). Export triggers gate scan if not already sanitized.

POST /api/verdicttank/rules (Feature 7, Enterprise/White-Label only)

{ "name": "SOC2 disclosure required", "template_id": "uuid",
  "rule_definition": { "condition": {"if": "vertical == 'fintech'"}, "required_fields": ["compliance.soc2_status"], "severity": "block" } }

Response 201 on pass, 422 with schema violation details on rule-syntax failure.

GET /api/verdicttank/rules / PATCH /api/verdicttank/rules/{id} / DELETE /api/verdicttank/rules/{id} - standard CRUD, org-scoped.

POST /api/verdicttank/org/branding (Feature 8, White-Label)

{ "logo_url": "...", "brand_colors": {"primary": "#0B1B33"}, "email_template": "<html>...</html>" }

POST /api/verdicttank/org/domain - initiates CNAME/TXT verification (see Section 5.6).

POST /api/verdicttank/reviews/{id}/peer-review (Feature 9)

{ "type": "roast", "comment": "Your CAC assumption ignores paid acquisition entirely." }

Rate-limited to 3/user/day (enforced via peer_review_rate_limits, Redis-cached counter). Returns 429 past cap.

POST /api/verdicttank/reviews/{id}/peer-review/{peer_review_id}/report - abuse reporting, pushes into moderation-queue.

GET /api/verdicttank/reviews/{id}/remediation - standalone fetch of the Fix-It plan (also embedded in verdict JSON).

GET /api/verdicttank/reviews/{id}/audit - standalone fetch of audit findings (Enterprise tier).


4. Pipeline Flow

4.1 Updated 6-Stage Pipeline

INTAKE                                                                        EGRESS
  β”‚                                                                              β”‚
  β–Ό                                                                              β”‚
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”   β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”   β”‚
β”‚ File Upload  β”‚   β”‚ URL Extract  β”‚   β”‚ Chat-to-     β”‚   β”‚  API submission   β”‚   β”‚
β”‚ (existing)   β”‚   β”‚ (Crawl4AI)   β”‚   β”‚ Refine draft β”‚   β”‚  (Enterprise)      β”‚   β”‚
β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”˜   β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”˜   β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”˜   β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜   β”‚
       β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜               β”‚
                                     β”‚                                            β”‚
                        normalize to standard intake format                       β”‚
                                     β”‚                                            β”‚
                                     β–Ό                                            β”‚
                    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”                            β”‚
                    β”‚  [Rules Engine pre-check:       β”‚  DEFERRED - Post-Launch   β”‚
                    β”‚   DEFERRED, see Β§1.2]           β”‚  (see Β§1.2, Β§5.7)         β”‚
                    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜                            β”‚
                                     β”‚                                            β”‚
    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”      β”‚
    β”‚ PHASE 1: Research Agent                                              β”‚      β”‚
    β”‚   Live web verification + citation gathering                         β”‚      β”‚
    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜      β”‚
                                     β”‚                                            β”‚
    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”      β”‚
    β”‚ PHASE 2: Primary Reviewer                                             β”‚      β”‚
    β”‚   10-dim brutal critique, 1-10 per dim                                β”‚      β”‚
    β”‚   + score_type tag (idea|proposal) per dim, overlay-only [v3.6, Β§2.2] β”‚      β”‚
    β”‚   + per-dimension explanation generated inline  [v3.6, Feature 4]     β”‚      β”‚
    β”‚   [org rule dimensions: DEFERRED - see Β§1.2, Β§5.7]                    β”‚      β”‚
    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜      β”‚
                                     β”‚                                            β”‚
    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”      β”‚
    β”‚ PHASE 3: Validation Reviewer - challenges primary score               β”‚      β”‚
    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜      β”‚
                                     β”‚                                            β”‚
             β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”                    β”‚
             β–Ό                                               β–Ό                    β”‚
    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”                          β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”            β”‚
    β”‚ PHASE 4:          β”‚                          β”‚ PHASE 4:          β”‚            β”‚
    β”‚ Cross-Check A      β”‚  (parallel)              β”‚ Cross-Check B      β”‚            β”‚
    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜                          β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜            β”‚
              β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜                    β”‚
                                      β–Ό                                            β”‚
                    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”                           β”‚
                    β”‚ PHASE 5: Majority Verdict          β”‚                           β”‚
                    β”‚  aggregates scores, builds          β”‚                           β”‚
                    β”‚  idea_score + proposal_score        β”‚                           β”‚
                    β”‚  as dual_score_overlay [v3.6, Β§2.2]  β”‚                           β”‚
                    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜                          β”‚
                                      β”‚                                             β”‚
                    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”                          β”‚
                    β”‚ [PHASE 6: Audit Agent - DEFERRED,       β”‚                          β”‚
                    β”‚  Post-Launch, see Β§1.2 / Β§5.6]           β”‚                          β”‚
                    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜                          β”‚
                                      β”‚                                             β”‚
                    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”                          β”‚
                    β”‚ POST-PIPELINE: Fix-It Generator          β”‚                          β”‚
                    β”‚  (Pro+; lightweight summary on Free)     β”‚                          β”‚
                    β”‚  quality-scored, 3-tier rubric [v3.6, Β§5.5]β”‚                          β”‚
                    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜                          β”‚
                                      β”‚                                             β”‚
                    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”                          β”‚
                    β”‚ SANITIZATION GATE (extended)              β”‚                          β”‚
                    β”‚  scans verdict JSON, explanations,         β”‚                          β”‚
                    β”‚  remediation text                          β”‚                          β”‚
                    β”‚  (audit findings: N/A, Phase 6 deferred)   β”‚                          β”‚
                    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜                          β”‚
                                      β”‚                                             β”‚
                    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”                          β”‚
                    β”‚ DISCLAIMER INJECTION (post-sanitization) │──────────────────────┐
                    β”‚  "no professional advice" on every        β”‚                      β”‚
                    β”‚  scored/explained/remediated surface [Β§2.6]β”‚                      β”‚
                    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜                      β”‚
                                      β”‚                                          β”‚
                    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”                     β”‚
                    β”‚ PDF Generation + Corpus Write               β”‚β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                    β”‚  (dual-score overlay, org-segmented)         β”‚
                    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

4.2 Data Flow Notes


5. Component Architecture

5.1 Chat-to-Refine - "Un-Copilot" Coaching (Ghostwriting Guard)

Priority Fix #4 (Judge 3): The v3.6 review found the original prompt-level ghostwriting mitigation ("system prompt says don't ghostwrite, flag suspicious output for manual review") reactive and non-deterministic - a coaching model instructed not to ghostwrite can still ghostwrite a full paragraph, and "flag for manual spot-check" catches it after the user already received it. Judge 3's recommendation: build a deterministic, real-time output-length guard that rejects multi-paragraph output before it reaches the user, and lean into an explicit "un-copilot" brand position - VerdictTank coaches by asking better questions, it does not write your pitch for you, and it says so out loud.

Real-time output-length guard (deterministic, not prompt-level):

  1. Every coach turn response is validated before it is returned to the user - not sampled, not spot-checked, every single turn. The guard runs synchronously in the chat-api request path immediately after the coaching-model call returns and before the response is persisted to chat_messages or sent to the client.
  2. Rejection criteria (any one trips the guard):
  3. Response contains more than one paragraph (defined as: more than one block separated by a blank line, or more than ~4 sentences in a single block - sentence-count is the fallback heuristic for models that omit paragraph breaks).
  4. Response exceeds 280 characters (SMS-length cap - deliberately short; a real Socratic question does not need three sentences of setup).
  5. Response does not contain a ? (i.e., it isn't phrased as a question at all - a purely declarative multi-sentence "here's what your pitch should say" response fails this check even if it's under the length cap).
  6. Response contains listable/structured markers characteristic of a drafted document rather than a conversational reply: markdown headers (#), numbered lists with more than 2 items, or bullet lists with more than 2 items.
  7. On rejection, the system does not simply retry the same prompt. It re-prompts the coaching model with an explicit corrective instruction appended to the turn: "Your previous response was too long / not a question. Respond with exactly one Socratic question (under 280 characters) that helps the user find the gap themselves. Do not draft content for them." This re-prompt is capped at 2 retries; if the model still fails to produce a compliant response on the 3rd attempt, the system falls back to a pre-written generic Socratic prompt from a static bank (e.g. "What's the one number in here you're least sure of?") rather than ever showing the user a rejected response.
  8. Every guard trip is logged (chat_messages.guard_rejected BOOLEAN, chat_messages.guard_rejection_reason TEXT, chat_messages.retry_count INT) for the quarterly manual audit in item 6 below, and to catch coaching-model/prompt regressions early (a spike in guard-trip rate on a model update is itself an alert-worthy signal, not just an audit finding).
  9. The guard is enforced server-side only - never trust a client-side check for this; the rejection logic lives in chat-api, not the web frontend, so it can't be bypassed by a different client hitting the same endpoint.

"Un-copilot" brand positioning (Judge 3's recommendation, product + UX + copy, not just engineering):

5.2 URL Extractor

5.3 Dual Scorer

5.4 Explainer

5.5 Fix-It Generator - Quality-Scored Remediation

Priority Fix #7 (Fix-It Quality Loophole): The v3.6 review found the original remediation-tracking design ("re-review score deltas... detect superficial template-filling vs. genuine improvement") was aspirational but not actually specified - there was no evaluation step that distinguished a user who genuinely added a TAM calculation from a user who pasted a single sentence into the relevant section just to make the presence-check pass. Presence-checking ("did the field get filled in") is trivially gameable and undermines the entire retention-loop premise of Fix-It as a credible improvement signal, not a checkbox exercise.

3-Tier Fix Quality Rubric (evaluation, not presence-checking):

On resubmission (re-review of a previously scored document, or a follow-up chat/URL submission linked via reviews.parent_review_id), each action_item that was addressed is evaluated against the specific gap it named - not merely against whether the target section now contains text. Evaluation runs as an additional judge-model call scoped narrowly to comparing the before/after content for a given action_item, using the original cited_gaps and description as the rubric anchor.

Tier Score Definition Scoring Rule
Superficial 1 The relevant section changed, but the specific gap named in cited_gaps is still unaddressed - e.g. a TAM number was added but with no visible methodology, source, or calculation shown; or text was added that mentions the topic without supplying the missing data/evidence. Triggers when the evaluator can find no calculation, citation, data point, or structural change that actually closes the gap - text presence alone does not clear this tier.
Minimal 2 The gap is nominally addressed but shallow - e.g. a TAM figure is present with a one-line methodology note, but no bottom-up/top-down hybrid breakdown, no source citation, no sensitivity range. Passes a "did they try" bar but not a "would this survive investor scrutiny" bar. Triggers when the evaluator finds a genuine attempt that directly responds to the cited gap, but missing at least one of: methodology transparency, supporting data/citation, or the specific technique the original action_item.description recommended.
Substantive 3 The gap is closed with the rigor the original action item called for - e.g. TAM calculated via the recommended hybrid formula, source-cited, with a stated methodology and range. The addition would plausibly change an informed reader's assessment of that dimension. Triggers when the evaluator confirms the specific technique/data named in the action item is present and internally consistent with the rest of the submission (a number that contradicts other stated figures elsewhere in the document does not qualify, even if superficially "complete").
Not Attempted 0 No detectable change to the relevant section between submissions. Default when the diffed section is identical or near-identical to the original.

5.6 Audit Agent [DEFERRED - Post-Launch, see Β§1.2]

This component is designed but not built for the 8-week MVP. Per the review's Week 1 Cut List, the Audit Agent addresses "an enterprise trust problem, not a launch problem" - it is retained here in full so the design isn't lost, and is revisited once Enterprise-tier pilot demand justifies the build (see Β§1.2 Deferred Trigger table).

5.7 Rules Engine [DEFERRED - Post-Launch, see Β§1.2]

This component is designed but not built for the 8-week MVP. Per the review's Week 1 Cut List, the Rules Engine is a White-Label/Enterprise dependency with no MVP demand signal. Retained here as the intended post-launch build.

5.8 White-Label Engine [DEFERRED - Post-Launch, see Β§1.2]

This component is designed but not built for the 8-week MVP, and additionally gated on Minimum Viable Legal (MSA + Privacy Program) per the review's Legal Killers before any Enterprise/White-Label customer is onboarded - see Β§2.6 and Β§1.2.

5.9 Community Layer (Roast/Boost) [DEFERRED - Post-Launch, see Β§1.2]

This component is designed but not built for the 8-week MVP. Per Judge 4's Legal Killer #4 (defamation/moderation exposure), the moderation overhead and liability surface outweigh launch-stage value - retained here as the intended post-launch build once moderation tooling and Minimum Viable Legal are in place.

6. Infrastructure

6.1 Server Layout

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚  Edge tier (2x app3-class instances, HA pair)                      β”‚
β”‚  Caddy - TLS termination, org-domain routing, rate limiting         β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚  App tier (autoscaling pool, 3-6 instances)                         β”‚
β”‚  review-api Β· chat-api Β· rules-api Β· org-api Β· peer-review-api      β”‚
β”‚  chat-ws (sticky sessions via Redis pub/sub for horizontal scale)   β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚  Worker tier (autoscaling pool, scales with queue depth)            β”‚
β”‚  pipeline-workers (Phases 1-6) Β· url-extractor Β· fixit-generator     β”‚
β”‚  sanitize-scanner Β· pdf-generator Β· domain-verifier Β· cron           β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚  Data tier                                                          β”‚
β”‚  PostgreSQL 16 (primary + read replica) w/ pgvector extension       β”‚
β”‚  Redis 7 (queue + cache + rate limits + feature flags)              β”‚
β”‚  S3-compatible object store (Wasabi/S3) for PDFs, transcripts,        β”‚
β”‚  URL snapshots, uploaded source docs                                 β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

6.2 Docker Services

services:
  edge-caddy:            {image: caddy:2, ports: ["443:443"]}
  review-api:            {build: ./services/review-api}
  chat-api:               {build: ./services/chat-api}
  chat-ws:                {build: ./services/chat-ws}
  rules-api:              {build: ./services/rules-api}
  org-api:                {build: ./services/org-api}
  peer-review-api:        {build: ./services/peer-review-api}
  pipeline-worker:        {build: ./workers/pipeline, deploy: {replicas: 4}}
  url-extractor:          {build: ./workers/url-extractor}   # wraps Crawl4AI
  fixit-generator:        {build: ./workers/fixit-generator}
  audit-agent-worker:     {build: ./workers/audit-agent}
  sanitize-scanner:       {build: ./workers/sanitize-scanner}
  pdf-generator:          {build: ./workers/pdf-generator}
  domain-verifier:        {build: ./workers/domain-verifier}
  moderation-worker:      {build: ./workers/moderation}
  cron:                   {build: ./workers/cron}            # T+90/180/365 predictions
  postgres:               {image: pgvector/pgvector:pg16, volumes: ["pgdata:/var/lib/postgresql/data"]}
  redis:                  {image: redis:7-alpine}

6.3 Database Choices

6.4 Caching

6.5 Deployment Options - ITPP-INFRA Shared vs Dedicated

This section addresses the single most common enterprise buyer question: "where does this run?" VerdictTank is not an abstract cloud service - it runs on bare-metal infrastructure owned and operated by IT Pro Partner, with two deployment models to match the client's risk and isolation requirements.

Option A: ITPP-INFRA Shared Deployment (Default)

The solution runs on IT Pro Partner's existing production infrastructure (ITPP-INFRA), the same hardware and backup pipeline that powers ITPP's own operations.

Layer Hardware Details
Edge netcup RS 4000 (app3) Caddy TLS termination, org-domain routing, rate limiting
App + Worker netcup RS 4000 (app2, app3) Containerized services (Docker), autoscaling within the pool
Data netcup RS 4000 (app2) PostgreSQL 16 + pgvector, Redis 7
Object Storage Wasabi S3 (shared bucket) PDFs, chat exports, URL snapshots, uploaded docs
Backup ITPP backup pipeline Daily 2-5 AM S3 sync, versioning ON, 90-day retention; database WAL archiving every 15 minutes via hermes-live-sync β†’ s3://itpropartner-backups/verdicttank/

Best for: Pro tier, early Enterprise pilots, non-regulated clients. Cost-efficient - leverages existing capacity with no dedicated hardware overhead. All infrastructure is documented and auditable under ITPP's existing SOC 2-type operational controls.

Option B: Dedicated Deployment

The solution runs on dedicated netcup or Hetzner bare-metal instances provisioned specifically for the client, with a dedicated S3 bucket and isolated backup pipeline.

Layer Hardware Details
Edge Dedicated netcup RS 2000/4000 or Hetzner CPX Isolated Caddy instance, client-specific TLS
App + Worker Dedicated netcup RS 4000 or equivalent No shared compute with other VerdictTank tenants
Data Dedicated PostgreSQL + Redis on the same instance or separate, per client requirements pgvector included
Object Storage Dedicated Wasabi S3 bucket No cross-tenant object storage; bucket owned by client or ITPP per contract
Backup Dedicated backup pipeline Independent S3 bucket with versioning ON; dedicated cron schedule; restore testing included quarterly

Best for: Enterprise clients with regulatory requirements (HIPAA, ITAR, FedRAMP-adjacent), White-Label resellers who need full data isolation guarantees, or any client whose compliance framework requires dedicated infrastructure with no shared-tenancy risk. All infrastructure is managed by IT Pro Partner - the client never touches servers - but the hardware, storage, and backup pipeline are theirs alone.

Shared responsibility line: In both models, ITPP manages everything below the application layer (OS, Docker, database, backups, monitoring, patching, incident response). The client's only responsibility is user account management and content submitted to the platform. This is the same operational model ITPP uses for all managed infrastructure - the client gets the benefit of bare-metal performance without any of the operational burden.


7. Security Model

7.1 Authentication & Authorization

7.2 Sanitization Gate (Extended for v3.6)

The v3 gate blocked deploys unless a pre-deploy scan confirmed no internal vendor/model identity or architecture detail could leak into public reports, PDFs, or API responses. v3.6 quadruples the freeform-text surface area the gate must cover:

Surface v3 v3.6
Public share reports βœ… βœ…
API responses βœ… βœ…
PDF reports βœ… βœ…
Dimension explanations (dimension_explanations.explanation_text) - βœ… new
Remediation action items (remediation_plans.action_items[].description) - βœ… new
Chat transcripts on export (chat_messages.content) - βœ… new
Audit findings (audit_findings.description) - βœ… new

7.3 Prompt Injection Defense (New Surface: Rules Engine + URL Extraction)

Two new v3.6 features introduce attacker-controlled or third-party text into a model context, which is new relative to v3's fully-controlled prompt pipeline:

7.4 Multi-Tenant Isolation (White-Label)

7.5 Abuse & Rate Limiting


8. Cost Model

8.1 Per-Review Cost Breakdown (v3.6, Pro/Enterprise Tier)

Cost Component v3 Baseline v3.6 Delta Running Total
Base pipeline (Research + Primary + 3 cross-checks) $0.47 - $0.47
3 specialist panel additions (Reasoning-Verification, Execution-Feasibility, Market-Reality) $0.21 - $0.68
Market simulation engine $0.06 - $0.74
Vertical red-team pass $0.05 - $0.79
Corpus write, embedding, similarity search $0.02 (dual-score schema, same query cost) $0.81
Prediction tracking cron (amortized) $0.01 - $0.82
Infrastructure (sanitization gate, PDF automation, storage) $0.04 (extended gate coverage, same infra cost) $0.86
v3 fully-loaded cost/review $0.86
Dual scoring + per-dimension explanation generation - +$0.04 $0.90
Fix-It action plan generation (Pro+ only) - +$0.05 $0.95
Second-Opinion Audit Agent (Enterprise only, amortized across all reviews) - +$0.03 $0.98
v3.6 fully-loaded cost/review $0.98

Chat-to-Refine and URL-to-Review add ~$0.02-0.03/session, treated as funnel/acquisition cost (not review COGS) - consistent with the free-tier loss-leader model. Free-tier full reviews remain ~$0.07/review; the lightweight Free-tier fix-it summary adds under $0.01.

8.2 Margin Check by Tier

Tier Price Included Reviews/mo Blended Rev/Review Cost/Review Gross Margin
Free $0/mo 1 - (loss leader) ~$0.07 n/a
Pro $79/mo 20 $3.95 $0.98 ~75%
Enterprise $499/mo 100 $4.99 ~$1.20 (incl. corpus/API/audit-agent infra) ~76%
White-Label $1,999+/mo fair-use unlimited volume-dependent ~$1.20-1.35 ~72-79%

8.3 Cost Attribution by Component (Engineering View)

Component Model calls added per review Notes
Chat-to-Refine N (per chat turn, pre-review, not per review) billed as acquisition cost, not COGS
URL Extractor 0 model calls (extraction is deterministic scraping) + sanitization pass negligible marginal cost
Dual Scorer 0 (tagging happens inside existing Phase 2 call) zero marginal model cost
Explainer 0 (explanation text generated inline in Phase 2 call, longer output token count) cost shows up as slightly higher Phase 2 token spend, captured in the $0.04 line
Fix-It Generator 1 additional call, post-pipeline $0.05/review, Pro+ only
Audit Agent 1 additional call, Phase 6 $0.03/review amortized, full cost when isolated to Enterprise-only billing is higher per Enterprise review
Rules Engine 0 additional model calls (injected into existing Phase 2 prompt + deterministic pre-check) cost is engineering/validation overhead, not inference
White-Label Engine 0 model calls pure infra/config cost
Community Layer 0 model calls (human-generated) moderation queue has infra cost, not inference cost

9. Implementation Sequence

Mapped to the v3.6 5-phase roadmap from the business proposal.

Phase 0 - Foundation (Weeks 1-3)

Phase 1 - Coach & Intake (Weeks 4-6)

Phase 2 - Transparency & Depth (Weeks 7-12)

Phase 3 - Fix-It & Trust (Weeks 13-19)

Phase 4 - Distribution & GA (Weeks 20-25)


10. Failure Modes

Component Failure Mode Detection Behavior / Recovery
Chat-to-Refine Coach model produces full-paragraph ghostwritten content instead of coaching questions Output length/completeness heuristic flags session for spot-check; periodic manual transcript audit (quarterly minimum) Session flagged, not blocked in real time (would break UX); pattern-level drift triggers prompt-tuning pass, not per-session blocking
Chat-to-Refine chat-ws socket drops mid-conversation Client-side reconnect with session resume via session_id REST polling fallback (POST /chat/{id}/message) available if WebSocket infra is degraded
URL Extractor Target page is JS-rendered SPA Crawl4AI can't fully execute Extraction returns near-empty content below a minimum-length threshold 422 with reason code extraction_insufficient; user prompted to upload file instead
URL Extractor Extraction exceeds timeout (45s hard cap) Worker timeout 422 extraction_timeout; no partial pipeline entry, no quota consumed
URL Extractor Extracted content contains injection payload ("ignore previous instructions...") Sanitization/injection-detection pre-pass on extracted text Content flagged and either stripped of the payload or the review is rejected with 422 content_flagged, never silently passed through
Dual Scorer Judges disagree on score_type tagging for the same dimension Audit Agent's tagging-consistency check (Phase 6) Logged as tagging_inconsistency finding; Phase 5 aggregation falls back to the majority tag among judges, review still completes
Explainer Primary Reviewer omits explanation.text or cited_gaps for a dimension Schema validation on Phase 2 output (fail-closed) Phase 2 call is retried once with an explicit "missing required field" repair prompt; second failure escalates to human review queue rather than shipping an incomplete verdict
Explainer Explanation text leaks internal architecture/vendor detail Sanitization gate scan Verdict withheld from egress until re-generated and re-scanned; review shows status=held_for_review to the user, not a silent partial response
Fix-It Generator Remediation plan generation call fails or times out Worker-level retry (3x exponential backoff) On exhaustion, verdict ships without remediation; a background job retries generation and notifies the user when the plan is ready (does not block verdict delivery)
Fix-It Generator Action items reveal too much evaluation logic (scoring-weight leakage) Manual spot-check + pattern detection on re-review score deltas (superficial compliance vs. genuine improvement) Prompt-level fix: regenerate remediation guidance templates to describe gaps, not weights; feed signal into reviewer-accuracy scoring per the proposal's mitigation
Audit Agent Shadow-mode audit never files a finding across many reviews Reviewer-accuracy-style tracking on the audit agent itself (judges.is_audit_agent=true) Treated as a signal the audit agent may be too conservative or echoing the primary panel; triggers prompt/config review before going live for billing
Audit Agent Audit call fails/times out Worker retry, then graceful degradation Verdict ships without Phase 6 findings (audit.findings=[], status still done, not blocked) - audit is additive, never a hard gate on verdict delivery
Rules Engine Org submits malformed rule JSON JSON Schema validation against rule_templates.schema_json at save time 422 with specific schema violation path; rule never persisted as is_active
Rules Engine Hard-block rule incorrectly blocks a valid review (false positive) Org admin dashboard shows block reason + rule name Org admin can deactivate/edit the specific rule (org_rules.is_active=false) without needing engineering support
Rules Engine Rule text attempts prompt injection via crafted field values Constrained AST parsing + fixed template rendering (Section 7.3) Injection payload is inert - it's substituted into a whitelisted template slot, never concatenated as raw instruction text
White-Label Custom domain CNAME/TXT never verifies domain-verifier polling times out after 48h Org notified with specific DNS record diagnostics; domain stays unverified, org continues on default subdomain until resolved
White-Label Cross-tenant data leak (application bug omits org_id filter) RLS policy at database layer blocks the query regardless of application bug Query returns empty/denied rather than cross-tenant data; incident logged via audit_log, alerts on-call
White-Label Branding CSS injection via brand_colors/logo_url fields Input validation (hex color regex, URL allowlist for logo hosting) Malformed values rejected at org-api before persisting to org_configs
Community Layer Roast/Boost comment contains harassment/abuse User-driven abuse reporting (peer_review/{id}/report) β†’ moderation queue Comment flagged (moderation_status='pending_review'), hidden from submitter view pending moderator action; repeated flags trigger user-level cooldown
Community Layer Rate limit bypass attempt (rapid submissions) Redis sliding-window counter + Postgres durable count, double-enforced Requests past the 3/day cap return 429; discrepancy between Redis and Postgres counts triggers a reconciliation job, not silent over-limit acceptance
Community Layer Feature accidentally enabled by default for a tier Feature flag (roast_boost_enabled) is off by default at the flag-service level, not the application code level Flag flip is a config change, auditable and instantly reversible without a deploy
Sanitization Gate (all surfaces) Gate itself fails/errors on a scan Fail-closed design: no scan result = no egress Content held, status=held_for_review; alerts on-call immediately since this blocks all outbound traffic for that review, by design
Cross-cutting Any Phase 1-6 model call fails Existing v3 degraded-mode fallback Pipeline runs with fewer judges rather than failing outright (v3 policy, unchanged in v3.6); dual-score/explanation/audit outputs degrade gracefully to "insufficient panel data" rather than fabricating scores

Appendix: Traceability to v3.6 Proposal Sections

Architecture Section Proposal Reference
Β§5.1 Chat-to-Refine Proposal Β§05, Feature 1
Β§5.2 URL Extractor Proposal Β§05, Feature 2
Β§5.3 Dual Scorer Proposal Β§06, Feature 3
Β§5.4 Explainer Proposal Β§06, Feature 4
Β§5.5 Fix-It Generator Proposal Β§07, Feature 5
Β§5.6 Audit Agent Proposal Β§08, Feature 6
Β§5.7 Rules Engine Proposal Β§08, Feature 7
Β§5.8 White-Label Engine Proposal Β§09, Feature 8
Β§5.9 Community Layer Proposal Β§09, Feature 9
Β§9 Implementation Sequence Proposal Β§12, Phases 0-4
Β§8 Cost Model Proposal Β§14, Financial Model
Β§7.2 Sanitization Gate Proposal Β§16, Condition 1
Β§7.4 Multi-Tenant Isolation Proposal Β§16, Condition 4

End of document.