AI-powered government contract proposal platform - find, analyze, and win federal contracts. Scores a proposal against the RFP for compliance, then tells you whether it will win.
"Throw your RFP in the Tank. Out comes a winning response."
The United States federal government spent $793 billion on contracts in FY2025. Every single dollar of that required someone to write a proposal. That proposal funnel is the choke point for an entire industry - and the tools serving it are broken.
The enterprise players hide pricing behind a sales call and target mid-to-large primes. The cheap tools are search wrappers with a chatbot bolted on. The middle - the 200,000+ active small businesses, independent consultants, and emerging contractors who collectively chase $182 billion in set-aside contracts every year - gets nothing purpose-built. They get spreadsheets, prayer, and $500/hour consultants they cannot afford.
RFP Tank fills that gap with a sledgehammer.
RFP Tank is a SaaS platform purpose-built for government contractors who need to find, analyze, and respond to RFPs faster and more accurately than any tool on the market. It is not a search tool with AI added. It is not an enterprise platform with a self-service tier bolted on. It is a complete proposal pipeline - from proactive opportunity matching to guided response flow to compliance checking to past performance mining - built from the ground up for the contractor who needs to WIN, not just submit.
The differentiation is twofold. First, depth on traceability: a compliance matrix that maps every mandatory requirement to the exact paragraph that answers it, with a confidence score and quoted evidence. Second, an evaluation layer no competitor has: an evaluator simulation that scores the proposal the way the government will, using the RFP's own criteria and weights, so the contractor knows whether it will win, not just whether it complies.
Market Reality:
Revenue Potential (Year 1 Conservative):
The competitive comparison is not close. GovEagle targets mid-to-large primes - no public pricing, no self-service, enterprise sales cycle only. CLEATUS is a SAM.gov wrapper. GovDash competes on price, not quality. Nobody has built the full guided proposal workflow for self-service contractors. Nobody. Until now.
RFP Tank is an IT Pro Partner product. The infrastructure is already running. The SAM.gov API is free. The USASpending.gov data is public. The value is in the processing, the workflow, and the intelligence layer - and that is exactly what we are building.
The Ask: Domain (rfptank.com - confirmed available), 90 days of focused development, and the confidence to price like the best product in the market - because that is what it will be.
GovEagle's pricing is unpublished and getting in the door requires an enterprise sales cycle. The other 99% of government contractors - the small businesses, independent consultants, and emerging prime candidates chasing $182 billion in set-aside contracts - are stuck writing proposals in Microsoft Word with a SAM.gov tab open and a spreadsheet for their compliance matrix.
RFP Tank is GovEagle for the other 99%.
Upload any RFP. In 60 seconds, get a complete compliance matrix, requirement breakdown, and Section L/M alignment. Then work through a guided response wizard that checks every answer against every requirement in real time. Build your past performance library once, mine it forever. Track every bid. Learn what wins. Before you submit, run an evaluator simulation that scores the proposal the way the government will, so you know whether it will win, not just whether it complies.
This is not a feature gap between us and the competition. This is market failure at $793 billion scale - and RFP Tank is the correction.
Government contracting is a $793 billion market with a proposal problem. Every awarded contract requires a winning response to a solicitation. The average federal RFP runs 50-200 pages. It contains dozens of mandatory requirements, multiple evaluation criteria, a compliance matrix that must be manually built, and formatting rules that disqualify non-conforming submissions. A single RFP response takes 40-200 hours of work. A mid-sized contractor might pursue 20-30 per year.
Here is how contractors currently handle this:
| Current Method | Who Uses It | Time Cost | Dollar Cost | Pain Level |
|---|---|---|---|---|
| Microsoft Word + SAM.gov tab | Most small contractors | 60-200 hrs/RFP | $0 tool cost, $15K-$50K labor | Extreme |
| Proposal consultants | Mid-sized firms | 40-100 hrs/RFP | $500-$1,500/hr | High + expensive |
| GovEagle / enterprise tools | Mid-to-large primes (50+ employees) | 20-60 hrs/RFP | Unpublished (enterprise) | Low tool friction, high cost |
| CLEATUS / search wrappers | Individual contractors | 40-100 hrs/RFP | $39-$78/mo | Moderate discovery, no response help |
| Spreadsheet compliance matrix | Everyone | 5-20 hrs per RFP | $0 | Tedious, error-prone, manual |
The result: small contractors are fighting $793 billion in opportunity with tools designed for a different era. They lose not because their work is inferior - they lose because their proposals are non-compliant, incomplete, or poorly structured. The requirement was there. They missed it. Or they addressed it halfway. Or they forgot to map it to an evaluation criterion. The government evaluator marks them down. The work goes to someone with a better proposal team.
RFP Tank serves three distinct users, all underserved:
Persona 1: The Solo Contractor
Persona 2: The Small GovCon Firm
Persona 3: The Emerging Prime
The market is not missing a search tool. SAM.gov is free. The market is missing a proposal execution engine - a tool that takes you from "I found an opportunity" to "I submitted a winning, compliant response" with AI assistance at every step. No existing tool does this for self-service users at any price point. That is the gap. It is not subtle. It is not a feature gap. It is a complete market absence at the most critical point in the contractor workflow.
The federal government contract market is one of the largest and most consistent spending pools in the world. It does not contract in recessions. It does not go to zero. It carries a statutory 23% small business contracting goal, roughly $182 billion of the $793 billion total.
| Market Segment | Annual Spend | Relevance to RFP Tank |
|---|---|---|
| Federal contract spending (FY2025) | $793 billion | Every dollar flows through a proposal |
| DoD contracts alone | $398.2 billion (50.2%) | Largest single agency pool |
| Small business set-asides | $182 billion (23% statutory goal) | Primary target user base |
| SLED (State, Local, Education) | ~$1.6 trillion | 2x the federal market, more accessible |
| Cooperative purchasing market | $65+ billion | Accessible via GS schedules and co-ops |
| Contracts expiring in 18 months | $142 billion (85,247 contracts) | Recompete pipeline - highest-value targets |
TAM: Roughly 200,000 active small business government contractors (per the SAM analysis in 4.2). At a $200/month average for a dedicated proposal tool, the addressable market is $480 million ARR.
Capture target: Even capturing 2% of the $480M TAM is $9.6M ARR - a venture-scale business from a small customer base. The near-term target (619 accounts, 0.31% penetration in 4.3) is deliberately conservative against this ceiling.
The key insight: A 1% win-rate improvement on the small business set-aside portion alone ($182 billion) is worth $1.8 billion in additional contract awards to small businesses. If RFP Tank helps a contractor win one additional $500K contract per year, it has delivered 50x the value of a $599/month subscription. That is the unit economics story that makes this product a no-brainer purchase.
| User Type | Estimated Count | RFP Tank Relevant | Serviceable |
|---|---|---|---|
| Active small business contractors (solo, firms, and emerging primes) | 200,000 | 100% relevant | 200,000 |
| Total SAM | ~200,000 |
The 200,000 is an umbrella spanning independent consultants, 2-20 person firms, and emerging primes of up to 100 employees - all of which qualify as small businesses under SBA size standards, counted once to avoid double-counting. The 500,000+ registered SAM.gov entities is the top-of-funnel universe; roughly 200,000 are active pursuers.
| Year | Target Subscribers | ARR | Penetration of SAM |
|---|---|---|---|
| Year 1 | 619 | $1.28M | 0.31% |
| Year 2 | 2,000 | $5.1M | 1.00% |
| Year 3 | 6,000 | $16.2M | 3.00% |
Year 1 penetration of 0.31% requires finding 619 customers in a market of 200,000. That is not a growth problem. That is a targeting problem - and GovCon is one of the most targetable communities in B2B software (LinkedIn by NAICS code, GovCon forums, APEX Accelerators, 8(a) cohort lists).
The competitive comparison is not close. It is embarrassing for the incumbents:
| Feature | RFP Tank | GovEagle | GovDash | CLEATUS | GovSignals | Vultron | Rohirrim | Proc. Sciences |
|---|---|---|---|---|---|---|---|---|
| Self-service signup | YES | NO | NO | YES | NO | NO | NO | NO |
| Transparent pricing | YES | NO | NO | YES | NO | NO | NO | NO |
| Profile-driven RFP matching | YES | NO | NO | Partial | NO | NO | NO | NO |
| RFP shredder (60-sec compliance matrix) | YES | YES | YES | NO | YES | YES | YES | YES |
| Guided step-by-step response flow | YES | NO | NO | NO | NO | NO | NO | NO |
| Real-time compliance check per answer | YES | NO | NO | NO | NO | NO | NO | NO |
| Past performance vault + AI mining | YES | Enterprise only | NO | NO | NO | NO | NO | NO |
| Win/loss analytics | YES | NO | NO | NO | NO | NO | NO | NO |
| Solo/individual tier | YES | NO | NO | YES | NO | NO | NO | NO |
| Team collaboration | YES | YES | YES | NO | YES | YES | YES | YES |
| Serves solo + SMB + mid-market | YES | NO | Partial | Partial | NO | NO | NO | NO |
| Walk-up entry price | $79/mo | Enterprise | Custom | $39/mo | Enterprise | Custom | Custom | Custom |
| Team-level price | $599/mo | Unpublished | Custom | $78/mo | Custom | Custom | Custom | Custom |
Nobody does all of this. Nobody. Until now.
RFP Tank column reflects planned MVP capability at launch. Build status by feature is tracked in Section 5.
+------------------------------------------------------------------+ | RFP TANK PLATFORM | +------------------------------------------------------------------+ | | | DATA INGESTION LAYER | | +-----------+ +--------------+ +------------------+ | | | SAM.gov | | USASpending | | State Procurement| | | | API (free)| | API (free) | | Portals (scraped)| | | +-----------+ +--------------+ +------------------+ | | | | | | | +---------------+-------------------+ | | | | | AI PROCESSING PIPELINE | | +-----------------------------------------------------+ | | | Opportunity Scorer | RFP Shredder | Content Miner | | | | (profile match) | (60-second) | (past perf) | | | +-----------------------------------------------------+ | | | | | USER LAYER | | +------------+ +------------------+ +------------------+ | | | Smart | | Guided Response | | Past Performance | | | | Discovery | | Flow + Compliance| | Vault + Mining | | | | Dashboard | | Copilot | | | | | +------------+ +------------------+ +------------------+ | | | | | OUTPUT LAYER | | +----------+ +----------------+ +------------------------+ | | | Word/PDF | | Compliance | | Win/Loss | | | | Export | | Matrix (Excel) | | Analytics Dashboard | | | +----------+ +----------------+ +------------------------+ | +------------------------------------------------------------------+
Technology stack: Node.js/Python backend, React frontend, PostgreSQL for user data, vector database for past performance content retrieval, LLM API for compliance checking and content generation.
This section addresses the single most common enterprise buyer question: "where does this run?" RFP Tank is not an abstract cloud service - it runs on bare-metal infrastructure owned and operated by IT Pro Partner, with two deployment models to match the client's risk and isolation requirements.
The solution runs on IT Pro Partner's existing production infrastructure (ITPP-INFRA), the same hardware and backup pipeline that powers ITPP's own operations.
| Layer | Hardware | Details |
|---|---|---|
| Edge | netcup RS 4000 (app3) or Hetzner CPX (app1-bu) | Caddy TLS termination, domain routing, rate limiting |
| App | netcup RS 4000 (app2, app3) or Hetzner CPX21 | Containerized services (Docker), autoscaling within the pool |
| Data | netcup RS 4000 (app2) | PostgreSQL 16 for user data + vector DB for past performance embeddings |
| Object Storage | Wasabi S3 (shared bucket) | Proposal archives, uploaded RFPs, generated response documents |
| Backup | ITPP backup pipeline | Daily S3 sync, versioning ON, 90-day retention; database WAL archiving via hermes-live-sync |
Best for: Solo tier, Consultant tier, early Team tier pilots. Cost-efficient - leverages existing capacity with no dedicated hardware overhead. All infrastructure is documented and auditable under ITPP's existing operational controls.
The solution runs on US-owned, US-hosted dedicated infrastructure - AWS GovCloud, Azure Government, or a US-owned colo - provisioned specifically for the client, with a dedicated object-storage bucket and isolated backup pipeline. This tier exists specifically for workloads that cannot run on shared infrastructure: CUI, ITAR-controlled data, or covered defense information under DFARS 252.204-7012.
| Layer | Hardware | Details |
|---|---|---|
| Edge | Dedicated AWS GovCloud or Azure Government | Isolated reverse proxy, client-specific TLS |
| App | Dedicated US-owned compute (GovCloud / Azure Gov) | No shared compute with other RFP Tank tenants |
| Data | Dedicated PostgreSQL + vector DB on the same instance or separate, per client requirements | Client data never co-mingles with other tenants |
| Object Storage | Dedicated US-region object storage (Wasabi US or GovCloud S3) | No cross-tenant object storage; bucket owned by client or ITPP per contract |
| Backup | Dedicated backup pipeline | Independent S3 bucket with versioning ON; dedicated cron schedule; restore testing included quarterly |
Best for: Enterprise clients handling CUI or ITAR-controlled solicitations, prime contractors with data-residency or US-person requirements, or organizations whose compliance framework requires dedicated, US-owned infrastructure with no shared-tenancy risk. The shared tier (Option A) is explicitly scoped to unclassified / FOUO-and-below data; anything at or above CUI routes to this tier. All infrastructure is managed by IT Pro Partner - the client never touches servers - but the hardware, storage, and backup pipeline are theirs alone.
Shared responsibility line: In both models, ITPP manages everything below the application layer (OS, Docker, database, backups, monitoring, patching, incident response). The client's only responsibility is user account management, profile configuration, and content submitted to the platform. This is the same operational model ITPP uses for all managed infrastructure - the client gets the benefit of bare-metal performance without any of the operational burden.
Feature 1: Smart Discovery (SETTLED - core design)
Contractor builds a one-time profile: NAICS codes, set-aside eligibility (8(a), SDVOSB, WOSB, HUBZone), geographic preferences, contract size range, agency history. RFP Tank ingests SAM.gov daily and scores every new solicitation against the profile. User receives a morning digest: "3 new opportunities match your profile. 1 is a recompete at an agency where you have past performance. 1 closes in 14 days." Recompete detection flags expiring contracts the user has previously bid or won. When a tracked solicitation is amended, RFP Tank notifies the contractor with exactly what changed - a moved due date, a new clause, an added requirement - not just "an amendment was posted."
Feature 2: RFP Shredder (SETTLED - core design)
Paste a SAM.gov URL, upload a PDF, or drop in a Word document. In 60 seconds: extracts all requirements, maps them to Section L (instructions) and Section M (evaluation criteria), identifies mandatory vs. desirable criteria, flags ambiguous language ("will be evaluated favorably" - what does that mean?), estimates competition level based on set-aside type and historical award data.
Same input path for the contractor's own proposal: upload the response document or paste its URL, and RFP Tank scores it against the RFP. This is the two-input model - RFP in, proposal in, compliance score out. Proposals are often multi-file (technical approach, cost volume, attachments), so ingestion accepts 1..N documents as one submission, not a single file.
Output: a structured compliance matrix. Every requirement, its source section, its evaluation weight, and a response field. This alone saves 5-20 hours per RFP.
Feature 3: Guided Response Flow (OPEN - primary differentiator)
NOT "here is an AI draft, good luck." A step-by-step wizard that walks the contractor through every requirement, section by section. For each requirement: write or paste your response, hit Review, and the AI checks: Does this address the requirement? Does it use government-recognized terminology? Is there quantifiable evidence? Does it map to the evaluation criterion? Green/amber/red per requirement. Progress tracker: "You have addressed 23 of 47 requirements. 3 are partially addressed. 2 are at risk."
This is the product nobody else has built. This is the feature that justifies the price.
Feature 4: Compliance Copilot (OPEN - builds on Feature 3)
Real-time AI assistant that watches as you type. Not post-hoc review - live checking. "This paragraph addresses Requirement 4.2 but does not quantify the outcome. Add a metric (e.g., reduced processing time by 40%)." Tracks every requirement across the full response. Nothing slips through.
Feature 5: Past Performance Vault (OPEN - Team+ tier)
A structured library of every past contract, capability statement, and relevant project description. AI mining: when a new RFP comes in, the vault surfaces relevant past performance automatically. "For this IDIQ requirement around IT modernization, you have 3 relevant past performance references. Click to insert and customize." Over time, the vault becomes the contractor's most valuable asset. It gets smarter with every proposal.
Feature 6: Win/Loss Intelligence (OPEN - Team+ tier)
Track every bid submitted through the platform. Mark wins and losses as awards are announced. Over time, the AI identifies patterns: win rate by agency, by NAICS, by contract size, by proposal section quality score. "Your win rate on DoD contracts is 42%. On DHS contracts it is 67%. Your strongest proposals score 85%+ on the Technical Approach section. Your weakest score below 60% on Management Approach." That insight is worth more than the subscription cost.
Feature 7: One-Click Export (SETTLED - standard)
Export the complete proposal in Word format (formatted per RFP instructions), PDF, or SharePoint-compatible package. Compliance matrix exports to Excel. Auto-generated cover letter. Everything formatted, nothing left to clean up.
The seven core features are the workflow. This layer is the moat. The gold standard is not a longer feature list - it is depth on the one thing every competitor does shallowly (traceability), plus an evaluation layer none of them do at all.
The three-product handoff keeps lanes clean: RFP Tank owns everything UP TO writing (extract, map, gap-check, score). ProposalTank writes the prose. VerdictTank reviews the finished draft. No feature bleed between them.
A. Traceability (the moat)
B. Evaluation (the differentiator, leverages the VerdictTank panel)
C. Mechanical compliance (boring but disqualifying)
An automated gate for page count, font, margins, per-volume page limits, required forms and signatures, and naming conventions. The top reason proposals are rejected unread.
D. Make-writing-a-breeze (still RFP Tank lane, no prose)
Prioritized build wedge
| Component | Status | Effort | Notes |
|---|---|---|---|
| SAM.gov API integration | TO BUILD | Low | Free API, well-documented, standard REST |
| USASpending.gov integration | TO BUILD | Low | Free API, historical award data |
| Opportunity scoring engine | TO BUILD | Medium | Profile-match algorithm, proprietary logic |
| RFP document parser (PDF/Word) | TO BUILD | Medium | Existing OSS libraries available |
| Compliance matrix generator | TO BUILD | Medium-High | Core IP, requires FAR/RFP domain logic |
| Guided response wizard | TO BUILD | High | Primary differentiator, custom build required |
| Compliance Copilot (real-time AI) | TO BUILD | High | LLM integration + real-time feedback loop |
| Past performance vault | TO BUILD | Medium | Document storage + vector retrieval |
| Win/loss analytics | TO BUILD | Medium | Database + dashboard, standard BI patterns |
| Word/PDF/Excel export | TO BUILD | Low-Medium | Existing OSS libraries available |
| Authentication + billing | TO BUILD | Low | Auth0 + Stripe, commodity |
| Marketing site (rfptank.com) | TO BUILD | Low | Standard ITPP deployment |
Infrastructure: Runs on existing ITPP app servers (app1/app3). No new servers required at MVP. LLM API costs are variable and passed through via usage-aware tier limits.
| Site | Status | Detail |
|---|---|---|
| rfptank.com | OPEN - not registered | Domain confirmed available, needs registration |
| app.rfptank.com | OPEN - not built | Main SaaS application |
| api.rfptank.com | OPEN - not built | Backend API |
| rfptank.com (marketing) | OPEN - not built | Marketing and signup site |
| Tier | Price | Users | Core Features |
|---|---|---|---|
| Solo | $79/mo ($948/yr) | 1 | Smart Discovery, RFP Shredder, Guided Response Flow, 5 active RFPs/mo |
| Consultant | $249/mo ($2,988/yr) | 1 | Everything in Solo + Compliance Copilot, Past Performance Vault (25 entries), Win/Loss tracking, 20 active RFPs/mo |
| Team | $599/mo ($7,188/yr) | Up to 5 | Everything in Consultant + multi-user collaboration, unlimited Past Performance Vault, advanced Win/Loss analytics, priority support, 50 active RFPs/mo |
| Enterprise | Custom (starts ~$2,000/mo) | Unlimited | Everything in Team + white-label, API access, dedicated CSM, custom NAICS/agency configurations, SLA, SSO |
Pricing rationale: GovEagle does not publish pricing and serves mid-to-large primes with a full sales cycle. We charge $79/month and serve the other 99% - walk-up, self-service, no demo required. The Solo tier is priced to convert on impulse. The Consultant tier is priced where a single proposal win pays for 5 years of the subscription. The Team tier is priced transparently at $599/month for up to five users, with no demo or sales call required.
| Stream | Source | Year 1 Revenue | Notes |
|---|---|---|---|
| Solo subscriptions | $79/mo per user | $148,520 | 157 subscribers avg |
| Consultant subscriptions | $249/mo per user | $185,505 | 62 subscribers avg |
| Team subscriptions | $599/mo per account | $138,369 | 19 accounts avg |
| Total Year 1 Revenue | $472,394 | Conservative ramp (619 accounts at Month 12) | |
| End-of-year ARR run rate | $1,278,972 | Month 12 MRR $106,581 x 12 |
Enterprise contracts (custom, ~$2,000/mo) are excluded from the base model; 2-3 added in Q4 would push Year 1 total above $500K.
This is not budget pricing. It is premium pricing positioned below enterprise alternatives. Here is the ROI math for each tier:
Solo at $79/month:
A solo contractor pursuing 10 RFPs per year spends 40-100 hours per response. At a blended rate of $100/hour in opportunity cost, that is $40,000-$100,000 in annual labor. Saving 30% of that is $12,000-$30,000 in value. The subscription costs $948/year. ROI is 12:1 to 30:1. This is a no-brainer at double the price. We price it at $79 to make the conversion decision require zero deliberation.
Consultant at $249/month:
A GovCon consultant who wins a single $300K task order earns a fee of $30K-$60K. If RFP Tank helps them win one additional contract per year that they would have otherwise lost, the subscription ($2,988/year) is irrelevant noise against a $30K fee. This tier also unlocks the Past Performance Vault - which becomes the consultant's competitive moat over time. Priced at $249 for transparent self-service access to an enterprise-grade feature set that GovEagle gates behind an unpublished enterprise sales process.
Team at $599/month:
A 5-person BD team at a small GovCon firm costs $300K-$600K in annual salaries. A 10% improvement in proposal quality that lifts win rate from 30% to 33% on a $10M pipeline generates $300K in additional contract awards. The subscription costs $7,188/year. The comparison is not close. We price at $599 because that is where a decision-maker can approve it without a committee.
Enterprise at Custom:
Enterprise accounts start at approximately $2,000/month and are scoped based on user count, API access, white-label requirements, and SLA needs. This is where the GovEagle comparison lands - we offer a comparable enterprise feature set with transparent pricing and self-service onboarding, versus GovEagle's unpublished enterprise pricing and 1-week implementation requirement.
| Cost Item | Monthly | Annual | Notes |
|---|---|---|---|
| LLM API (OpenAI/Anthropic) | $800-$2,500 | $9,600-$30,000 | Scales with usage; tier limits control cost |
| Server infrastructure | $0 | $0 | Already running on ITPP infra |
| SAM.gov API | $0 | $0 | Free public API |
| USASpending.gov API | $0 | $0 | Free public API |
| Auth0 (authentication) | $70 | $840 | Up to 7,000 MAU on free tier initially |
| Stripe (payment processing) | 2.9% + $0.30/transaction | ~$2,000-$8,000 | Variable with revenue |
| Domain + SSL | $15 | $180 | rfptank.com registration |
| Monitoring + logging | $50 | $600 | Datadog or equivalent |
| Customer support tooling | $100 | $1,200 | Help desk software |
| Marketing (content + ads) | $2,000 | $24,000 | Primarily content + LinkedIn |
| Total Monthly OpEx | ~$3,000-$5,000 (early months) | ~$36,000-$60,000 | Ramps to ~$11,000/month at scale (see 10.1) |
Key insight: The infrastructure advantage is real. Other startups in this space are paying $5,000-$15,000/month in AWS/GCP costs before they write a line of product code. ITPP infrastructure is already running. That is a structural cost advantage that compounds as we scale.
At full scale (619+ accounts), estimated gross margin is 78-85%. The primary variable cost is LLM API usage, which is manageable through tier-based usage limits (RFPs per month, compliance checks per day). A single compliance check is 1-2 LLM calls of roughly 1,000-2,000 tokens, or about $0.02-$0.05 at current Claude/GPT pricing; a heavy session of 50 checks costs $1.00-$2.50, leaving the $79/month Solo tier at 80%+ gross margin even under daily-use limits. Real-time copilot latency is bounded to roughly 1-2 seconds per check with streaming, acceptable for an inline writing assistant. Unlike a services business, the marginal cost of serving the 501st customer is near zero.
| Month | Solo Subs | Consultant Subs | Team Accounts | MRR | Cumulative Revenue |
|---|---|---|---|---|---|
| 1 | 10 | 2 | 0 | $1,288 | $1,288 |
| 2 | 20 | 5 | 1 | $3,424 | $4,712 |
| 3 | 35 | 10 | 2 | $6,453 | $11,165 |
| 4 | 55 | 18 | 4 | $11,223 | $22,388 |
| 5 | 80 | 28 | 7 | $17,485 | $39,873 |
| 6 | 110 | 40 | 11 | $25,239 | $65,112 |
| 7 | 145 | 55 | 16 | $34,734 | $99,846 |
| 8 | 185 | 72 | 22 | $45,721 | $145,567 |
| 9 | 230 | 92 | 29 | $58,449 | $204,016 |
| 10 | 280 | 115 | 37 | $72,918 | $276,934 |
| 11 | 335 | 140 | 46 | $88,879 | $365,813 |
| 12 | 395 | 168 | 56 | $106,581 | $472,394 |
Year 1 total revenue: ~$472K (conservative ramp, 0% churn assumption in early months is unrealistic - see Risk section).
End-of-year MRR run rate: $106,581 ($1.28M ARR) across 619 paying accounts.
These numbers assume no enterprise contracts in Year 1. With 2-3 enterprise accounts added in Q4, Year 1 total exceeds $500K and end-of-year ARR exceeds $1.3M.
These are not features. They are structural advantages that compound over time. Each one is hard to copy individually. Together, they are a fortress.
To be precise about what is and is not durable: three of the seven are structural advantages that take real build time and GovCon domain knowledge (profile-driven matching, guided response flow, real-time compliance). Two are positioning choices a competitor could copy in a quarter if it decided to (self-service pricing, full-stack coverage). The durable, compounding moat is Knockout 4 plus Section 7.3 - the Past Performance Vault and win/loss data lock-in, which cannot be copied because the value lives in the user's own data, not in our code.
Knockout 1: Profile-Driven Proactive Matching
Every other tool is reactive - you go find the RFP. RFP Tank finds the RFP for you, scores it against your profile, and tells you which ones are worth your time before you ever open the solicitation. Nobody else does this end-to-end. It requires building a profile engine, a scoring algorithm, and a daily ingestion pipeline. CLEATUS does partial matching but without the full profile context or the scoring depth.
Knockout 2: Guided Response Flow
This is the product nobody built. It is the feature that should exist and does not. The reason: building it correctly requires understanding how government proposals are evaluated - Section L/M alignment, FAR clause awareness, evaluation factor weighting. That knowledge is not in engineering departments. It is in proposal professionals. We are building it first, ahead of everything else, because it is the differentiator the rest of the product hangs on. No competitor has it, and none is positioned to build it quickly.
Knockout 3: Real-Time Compliance Copilot
Every other AI in this space does post-hoc review: "Here is a score for your completed proposal." We check in real time, as you write. The distinction is not cosmetic - catching a compliance gap while writing is 10x cheaper than catching it after you have written 40 pages that need to be rewritten. This is a workflow advantage that shows up in win rates, not just user experience.
Knockout 4: Past Performance Vault With AI Mining
Your past performance is your biggest competitive asset in government contracting. It is also scattered across email chains, old Word documents, and the memory of your BD team. The Vault organizes it, and the AI mines it automatically for each new opportunity. Over time, users who have populated their vault have a compounding advantage over users who have not - and over competitors whose tools do not offer this capability at all.
Knockout 5: Win/Loss Analytics
This is CRM for proposal professionals. Track every bid, mark outcomes, and let the AI identify what separates your wins from your losses. No competitor offers this at the SMB level. Enterprise BD teams at large primes do this manually with spreadsheets. We automate it for solo contractors and 5-person BD teams at a fraction of the cost.
Knockout 6: Transparent Self-Service Pricing
The GovCon software market is dominated by "book a demo" gatekeeping. GovEagle, GovDash, Vultron, Rohirrim - none of them will tell you what they charge without a sales call. That is a conversion killer for small contractors who do not have time for a 3-meeting sales cycle. RFP Tank has a pricing page. You can be live in 5 minutes. No demo. No call. No contract negotiation. That is a product-led growth strategy that scales without a sales team.
Knockout 7: Full-Stack Workflow for Any Contractor Size
Every competitor serves a slice of the market. GovEagle serves large primes. CLEATUS serves individuals (with a shallow product). Nobody serves the full spectrum - from solo consultant to 50-person firm - with a single product that scales appropriately. RFP Tank does. One platform, four tiers, zero workflow gaps. A solo contractor who grows to a 10-person firm does not need to switch tools. They upgrade their tier.
HIGH CAPABILITY
|
RFP Tank | GovEagle
[Solo-Ent] | [Enterprise]
|
SELF-SERVICE ---------+--------- ENTERPRISE-ONLY
$79/mo | Unpublished
|
CLEATUS | GovDash
[Search] | [Price-led]
|
LOW CAPABILITY
RFP Tank occupies the upper-left quadrant - high capability, self-service access. No competitor is positioned there. That is not an accident. Building high capability AND self-service simultaneously is hard. It requires product discipline that most enterprise-focused companies cannot apply because their incentive is to serve enterprise customers with white-glove support, not to make the product simple enough for a solo contractor to onboard in 5 minutes.
The Past Performance Vault creates data lock-in. Users who have spent 6 months populating their vault with past contracts, capability statements, and proposal content are not switching tools. The switching cost is not financial - it is the loss of an organized, AI-indexed library that took months to build. This is the same moat Salesforce uses (CRM data lock-in), applied to the GovCon proposal workflow.
Win/Loss data deepens the moat further. After 12 months of tracking bids and outcomes, the platform knows which proposal sections correlate with wins for that specific contractor, at that specific agency, in that specific NAICS. That intelligence is contractor-specific and cannot be transferred to a competitor's tool.
| Phase | Timeline | Goal | Key Activities |
|---|---|---|---|
| Foundation | Months 1-2 | Build the product + pre-launch waitlist | Core feature development, rfptank.com marketing site live, LinkedIn presence, waitlist at 200+ |
| Soft Launch | Month 3 | First 50 paying customers | Invite waitlist, APEX Accelerator outreach, GovCon LinkedIn content, beta feedback loop |
| Growth | Months 4-8 | Reach 200 paying customers | Content marketing, GovCon forum presence, referral program, first case studies |
| Scale | Months 9-12 | 500+ subscribers, first enterprise accounts | Paid LinkedIn ads, APMP partnerships, conference presence, enterprise sales motion |
| Channel | Tactic | Cost | Expected Reach | Expected Conversion |
|---|---|---|---|---|
| LinkedIn organic | Weekly GovCon content, founder posts about the product | $0 | 2,000-10,000 per post | 1-3% to signup |
| APEX Accelerator partnerships | Email outreach to 50 regional APEX Accelerators | $0 | 50-200 referrals per active APEX Accelerator | High - warm referrals |
| LinkedIn paid ads | Target: NAICS code keywords, 8(a) designation, SAM.gov users | $2,000-$5,000/mo | 50,000-100,000 impressions | 0.5-1% CTR, 10% trial conversion |
| Content / SEO | RFP compliance, proposal writing how-to articles | $500/mo (writer) | Compound - 5,000-20,000 visits/mo by Month 12 | 2-4% trial signup |
| GovCon forums | Authentic community participation + tool mentions | $0 | 500-2,000 per active thread | 5-10% (trust-based) |
| APMP membership | Attend chapter meetings, sponsor newsletter | $1,000-$3,000/yr | 5,000 proposal professionals | 2-5% trial signup |
| Referral program | 2 months free for each paying referral | Revenue share | Compounds with customer base | 30-50% of growth by Month 9 |
| Cold outreach | LinkedIn DM to 8(a) and SDVOSB firm principals | $0 | 100 contacts/week | 5-10% response, 2-3% trial |
All conversion rates above are planning hypotheses based on benchmark SaaS funnels and adjacent GovCon communities; none have been field-tested. The first 90 days are scoped to validate these against real channel data before scaling spend (see 9.6).
The RFP Shredder is a natural free-trial hook. Let any user - without an account - paste a SAM.gov opportunity link and get a preview compliance matrix (first 10 requirements extracted, rest gated). This creates an immediate "holy sh*t" moment. The prospect sees the value before they give us an email address. When they hit the gate on requirement #11, they sign up. This is how Figma, Loom, and Notion grew their initial user bases - lead with the product, not the pitch.
| Risk | Probability | Impact | Mitigation |
|---|---|---|---|
| Federal contract spending cuts (budget sequestration) | Low | High | SLED market ($1.6T) is not federally funded; pivot emphasis to state/local if federal spend contracts significantly |
| Government moves RFP process fully to AI-generated solicitations, reducing proposal complexity | Very Low | High | This is 10-15 years away at minimum; FAR reform is glacial |
| SAM.gov API changes or rate limiting | Medium | Medium | Build abstraction layer; maintain cached opportunity database; FedBizOpps has multiple unofficial aggregators as backup |
| GovCon market contraction (agency consolidations) | Low | Medium | More consolidation typically means larger contracts with more competitive bidding - a tailwind for proposal tools |
| Small business set-aside percentages reduced by policy change | Very Low | Medium | Total contract spend growing regardless; larger share of smaller pie still represents accessible market |
| Risk | Probability | Impact | Mitigation |
|---|---|---|---|
| Compliance matrix AI produces errors that cause proposal disqualification | Medium | High | Build in explicit disclaimer; require user review of all AI output; red-flag high-stakes requirements for manual verification; never position AI as a replacement for expert review |
| LLM API costs exceed projections as usage scales | Medium | Medium | Implement strict per-tier usage limits from Day 1; build caching for common document types; negotiate enterprise API pricing above 10M tokens/month |
| 60-second RFP shredder cannot handle all document formats (scanned PDFs, unusual formatting) | High | Low-Medium | Build graceful degradation - partial extraction is better than failure; queue for manual review on complex documents |
| Guided response flow UX is too complex for solo contractors | Medium | High | Conduct user testing with 10-20 target users before launch; iterate on flow based on completion rates; provide skip/simplify options |
| Past performance vault has slow adoption (users do not populate it) | High | Low | Build import tools (upload old proposals, pull from USASpending); seed with public data; gamify completion |
| Risk | Probability | Impact | Mitigation |
|---|---|---|---|
| GovEagle launches a self-service tier | Low | High | Their sales model, org structure, and positioning make this nearly impossible to execute without a complete company pivot; we would still have 12-18 months of market lead |
| CLEATUS adds a guided response flow | Medium | Medium | CLEATUS is a data company - adding a workflow layer to a search tool requires rebuilding the product; their feature velocity is slow |
| New entrant raises $10M+ to build exactly what we are building | Medium | High | Speed is the mitigation; launch in 90 days, get to 100 customers before they get to product-market fit; first-mover + data moat compounds quickly |
| OpenAI or similar releases a general-purpose "RFP assistant" | Medium | Low-Medium | Generic tools cannot replicate the SAM.gov integration, profile matching, and FAR-specific compliance logic; government contracting is too specialized for a general assistant |
| GovDash copies the guided response flow | Low | Medium | Their positioning is "cheap alternative" - adding a premium workflow feature contradicts their own marketing |
| Risk | Probability | Impact | Mitigation |
|---|---|---|---|
| Single-developer bottleneck (Germaine) | High | High | Document architecture thoroughly from Day 1; pre-commit a contract developer for the SAM.gov ingestion pipeline and compliance-matrix backend from Week 4 (fixed ~$2,000-$4,000, not gated on traction); open source non-core components to attract contributors |
| Customer support load overwhelms available capacity | Medium | Medium | Invest in self-service documentation from launch; build FAQ and tutorial library before the first 50 customers; implement ticket triage by tier |
| Data breach or security incident exposes contractor proposal data | Low | Very High | Contractor proposals contain competitively sensitive pricing, teaming information, and proprietary technical approaches; encrypt all data at rest and in transit; implement SOC 2 Type II audit by Year 2; no CUI handling on the shared tier (FedRAMP not required there); CUI and ITAR workloads route to the US-owned dedicated tier (Option B), built to DFARS 252.204-7012 / FedRAMP Moderate-equivalent standards |
| Key customer (enterprise) churns in first 6 months | Medium | Medium | Enterprise customer success process from Day 1; dedicated Slack channel for enterprise accounts; monthly check-ins; usage monitoring to catch disengagement early |
| No assumption validated with a customer yet | High | High | Every number in this document (willingness-to-pay, conversion rates, CAC) is a hypothesis until tested; run 10 discovery interviews and a waitlist before scaling marketing spend; kill or reprice tiers the data does not support |
The most likely failure modes, in order of probability:
1. The guided response flow ships too slow. If the differentiating feature - the guided wizard with real-time compliance checking - takes 6 months to build, we launch with just the RFP Shredder. The Shredder is useful, but it is not a sufficient moat. GovDash and GovEagle both have compliance matrix generation. Without the guided flow, we are just another RFP tool. Mitigation: the guided flow is feature #1, not feature #7. Build it first, ship everything else after.
2. Trial-to-paid conversion is below 5%. If users love the free RFP Shredder preview but do not convert to paying subscribers, the product-led growth loop breaks. This usually means the paywall hits before users see enough value, or the paid features are not differentiated enough from the preview. Mitigation: track conversion rates weekly from Day 1; iterate the paywall and trial experience aggressively in the first 60 days.
3. LLM API costs eat the margin. At scale, if users are running dozens of compliance checks per session and the per-check LLM cost is not bounded, gross margin compresses to zero. Mitigation: enforce per-tier daily limits from Day 1; cache common document parsing results; monitor API spend daily.
4. GovCon community does not trust a new tool with proposal data. Government contractors are paranoid about data security for good reason - proposals contain pricing strategies, teaming agreements, and proprietary technical approaches. If RFP Tank cannot answer basic security questions (encryption, data retention, who can access my data), the trust barrier kills conversion. Mitigation: publish a clear security page before launch; answer these questions proactively on the pricing page.
5. Product-market fit is in Consultant tier, not Solo. The economics of Solo ($79/month) require high volume to generate meaningful revenue. If product-market fit turns out to live in the Consultant ($249) and Team ($599) tiers - which is plausible, since those users have more acute pain and more financial stake in win rates - the marketing strategy needs to tilt toward that segment immediately. Mitigation: track ARPU and NPS by tier; respond to where the enthusiasts actually are.
| Metric | Target | Failure Signal |
|---|---|---|
| Waitlist signups before launch | 500+ | Under 200 = insufficient demand signal |
| Trial signups in first 30 days post-launch | 100+ | Under 50 = distribution problem |
| Trial-to-paid conversion | 20%+ | Under 10% = pricing or value proposition problem |
| Month 3 paying customers | 50+ | Under 25 = growth rate will not hit Year 1 target |
| Net Promoter Score (NPS) | 40+ | Under 25 = product is not delivering on promise |
| Churn in first 90 days | Under 9% | Over 15% = product-market fit problem |
| Average session length | 20+ minutes | Under 10 minutes = users are not engaging with the workflow |
| Month | Revenue | LLM API | Marketing | Labor (est.) | Other OpEx | Net Income |
|---|---|---|---|---|---|---|
| 1 | $1,288 | $200 | $1,000 | $5,000 | $500 | -$5,412 |
| 2 | $3,424 | $400 | $1,500 | $5,000 | $500 | -$3,976 |
| 3 | $6,453 | $600 | $2,000 | $5,000 | $600 | -$1,747 |
| 4 | $11,223 | $900 | $2,500 | $5,000 | $700 | $2,123 |
| 5 | $17,485 | $1,200 | $3,000 | $5,000 | $800 | $7,485 |
| 6 | $25,239 | $1,600 | $3,500 | $5,000 | $900 | $14,239 |
| 7 | $34,734 | $2,000 | $4,000 | $5,000 | $1,000 | $22,734 |
| 8 | $45,721 | $2,400 | $4,500 | $5,000 | $1,000 | $32,821 |
| 9 | $58,449 | $3,000 | $5,000 | $7,500 | $1,200 | $41,749 |
| 10 | $72,918 | $3,500 | $5,000 | $7,500 | $1,200 | $55,718 |
| 11 | $88,879 | $4,000 | $5,000 | $7,500 | $1,500 | $70,879 |
| 12 | $106,581 | $4,500 | $5,000 | $7,500 | $1,500 | $88,081 |
| TOTAL | $472,394 | $24,300 | $42,000 | $70,000 | $11,400 | $324,694 |
Notes on assumptions:
| Metric | Solo | Consultant | Team | Enterprise |
|---|---|---|---|---|
| Monthly price | $79 | $249 | $599 | ~$2,000 |
| Gross margin (estimated) | 82% | 80% | 78% | 75% |
| Assumed monthly churn | 3% | 2% | 1.5% | 1% |
| Average customer lifetime | 33 months | 50 months | 67 months | 100 months |
| LTV (lifetime value, gross profit) | $2,138 | $9,960 | $31,304 | $150,000 |
| Estimated CAC (all channels) | $80 | $150 | $400 | $2,000 |
| LTV/CAC ratio | 27:1 | 66:1 | 78:1 | 75:1 |
LTV/CAC ratios above 3:1 are considered healthy for SaaS. Ratios above 10:1 indicate potential underinvestment in growth. All four tiers show ratios in the 27:1 to 78:1 range. Note these ratios rest on CAC figures that are planning estimates, not field-tested channel data; they indicate room to increase acquisition spend once real CAC is measured, not permission to spend ahead of validation. On churn: the 3% monthly Solo assumption compounds to ~8.7% over 90 days, just under the 'Under 9%' success threshold in section 9.6 - a deliberately tight margin that makes early churn the model's most sensitive assumption and the primary Month-4 go/no-go gate.
At the early-stage cost structure from 6.3 (~$3,000-$5,000/month in OpEx excluding labor), the business turns cash-flow positive in Month 2 and net-income positive (including a $5,000/month labor allocation for Germaine's opportunity cost) in Month 4, at 77 paying subscribers and ~$11,200 MRR.
The $10,000-$12,000/month OpEx figure is the Month 11-12 scale structure (see 10.1), reached only after the business is running ~$90,000/month in net income. Break-even is an early-months event, not a scale-stage milestone - a realistic and achievable timeline for a product with genuine product-market fit.
| Scenario | Subscriber Growth | Year 1 Revenue | End ARR |
|---|---|---|---|
| Conservative | 15%/mo from Month 3 | $472K | $1.28M |
| Realistic | 20%/mo from Month 3 + 3 enterprise | $520K | $1.45M |
| Aggressive | 30%/mo from Month 3 + 6 enterprise, partnership channels activate | $850K | $2.1M |
The conservative scenario requires 619 paying accounts by Month 12. The realistic scenario adds enterprise contracts and assumes one major APEX Accelerator or APMP channel activation. The aggressive scenario assumes a product that becomes the default recommendation in GovCon communities by Month 6 - plausible if the guided response flow delivers on its promise.
| Year | Subscribers | ARR | Gross Profit | Notes |
|---|---|---|---|---|
| Year 1 | 619 | $1.28M | $1.0M | Build + launch + first traction |
| Year 2 | 2,000 | $5.1M | $4.0M | Scale GTM, hire sales + support |
| Year 3 | 6,000 | $16.2M | $13.0M | Enterprise motion + market leadership |
Year 3 ARR of $16M at 80% gross margin is a $130M+ valuation at standard SaaS multiples (8-10x ARR). This is not a lifestyle business. It is a venture-scale outcome from a $0 infrastructure investment and a domain that costs $15/year.
| Item | Details | Timeline | Estimated Cost |
|---|---|---|---|
| Domain registration (rfptank.com) | Confirmed available - register immediately | Day 1 | $15/year |
| SAM.gov API key | Free public registration at api.sam.gov | Week 1 | $0 |
| USASpending.gov API key | Free public registration | Week 1 | $0 |
| LLM API access (OpenAI or Anthropic) | Production API key with usage monitoring | Week 1 | $50-$200 seed deposit |
| Stripe account setup | Payment processing for subscriptions | Week 1 | $0 upfront |
| Auth0 account setup | Authentication (free tier to 7,000 MAU) | Week 1 | $0 initially |
| Marketing site (rfptank.com) | Deploy on existing ITPP infrastructure | Weeks 1-2 | $0 |
| Customer discovery interviews | 10 GovCon contractor interviews validating willingness-to-pay and feature priorities | Weeks 1-3 | $0 |
| Core product development | SAM.gov ingestion + RFP Shredder + Guided Flow MVP | Months 1-2 | Germaine's time |
| Pre-launch LinkedIn content | 8-10 posts building waitlist | Months 1-2 | $0 |
| APEX Accelerator outreach campaign | Email to 20-30 regional APEX Accelerator offices | Month 2 | $0 |
| Beta program (10-20 users) | Invite from waitlist for pre-launch feedback | Month 3 | $0 or comped subscriptions |
| Contract development support | SAM.gov ingestion + compliance-matrix backend | Weeks 4-10 | $2,000-$4,000 |
| Launch marketing push | LinkedIn, content, community announcements | Month 3 | $500-$1,000 |
Total hard dollar cost to launch: ~$2,500-$5,000 (contract-developer contingency plus marketing), removing the single point of failure without waiting for traction. The rest is ITPP infrastructure (already paid for) and Germaine's time. This is the single best risk/reward ratio in the current ITPP product portfolio.
These decisions need answers before Week 2. Each one unblocks a subsequent action.
Decision 1: Domain - Register rfptank.com NOW
| Domain | Status | Verdict |
|---|---|---|
| rfptank.com | AVAILABLE - confirmed | Top pick - direct, memorable, product-literal. Register today. |
| rfptank.io | Check availability | Acceptable fallback if .com somehow expires before registration |
| rfpshredder.com | Not checked | Describes one feature, not the full product - weaker brand |
| govproposal.com | Not checked | Generic, not memorable |
| proposaltank.com | Not checked | Acceptable alternative if rfptank.com ever became unavailable |
Recommendation: Register rfptank.com today. Domains at this price point that are available do not stay available if the concept leaks. This is a $15 decision with potentially significant downstream value. There is no reason to delay.
Decision 2: Product vs. Consulting Play
RFP Tank is an IT Pro Partner SaaS product. It is not a GovCon consulting practice. The product helps others win contracts; ITPP does not bid contracts using it (yet). There is an optional future play where Germaine's veteran colleague and ITPP form a minority-veteran-owned JV that USES RFP Tank to bid on IT contracts - creating live case studies and a demonstration environment. But the product itself launches as ITPP software. Confirm: product play first, consulting JV optionally later.
Decision 3: Build Sequence
The guided response flow is the differentiator. It should be Feature 1, not Feature 7. Proposed sequence:
This sequence carries two weeks of built-in slack. The guided response wizard MVP (4 weeks) is the critical path, and the daily SAM.gov ingestion is contracted out (see 11.1) so it does not compete for Germaine's time. If the guided flow slips, beta still ships on the RFP Shredder + profile matching (already differentiated), and the guided flow lands in Week 12.
Confirm this build order before Day 1 of development.
Decision 4: Pricing Confirmation
The tiers proposed ($79/$249/$599/custom) are recommendations. They are priced to be premium without being enterprise. If Germaine wants to price higher (e.g., Solo at $99, Consultant at $299) the unit economics only improve. If the goal is maximum early traction, the current tiers are correct. Confirm pricing before the marketing site goes live.
Decision 5: IT Pro Partner Branding
Does rfptank.com show "Powered by IT Pro Partner" or does it stand alone as a brand? Most SaaS products benefit from standalone branding for credibility in a niche market. ITPP is an MSP brand; RFP Tank is a GovCon software brand. They serve different audiences. Recommendation: standalone brand for rfptank.com, with a discreet "A product of IT Pro Partner" attribution in the footer.
| Metric | Target |
|---|---|
| Paying accounts | 619+ |
| Monthly Recurring Revenue | $106,000+ |
| Annual Recurring Revenue | $1.28M+ |
| Net Promoter Score | 45+ |
| Churn rate (monthly) | Under 3% |
| Trial-to-paid conversion | 20%+ |
| Enterprise accounts | 2-5 |
| APEX Accelerator partnerships active | 5+ regional APEX Accelerators recommending RFP Tank |
| Case studies published | 3+ contractors with documented win-rate improvements |
The benchmark for "product-market fit confirmed" is Month 4: 50+ paying subscribers who were NOT given the product for free, with NPS above 40 and monthly churn below 3%. If those three numbers are green at Month 4, this product grows to $1.28M ARR. If they are red, we know what to fix before we have spent real marketing dollars.
RFP Tank is not just a product. It is a position in a $793 billion market that nobody has credibly claimed at the self-service level. The government contracting industry is underserved at the small business layer because the tools that exist were built for the large primes, not for the contractors who need them most.
Every month that passes without RFP Tank is a month that 50,000+ small business contractors submit proposals in Word, built with spreadsheet compliance matrices, without a single line of AI assistance on whether they actually addressed the requirement they think they addressed. Some of them will lose contracts they deserved to win. Some of them will exit the GovCon market entirely, convinced the system is rigged, when the real problem was a solvable workflow failure.
RFP Tank does not just serve the market. It corrects a market failure at scale. That is the kind of business worth building.
The competition saw what was possible in this market. They served the 1% and ignored the rest. That was their plan. And then RFP Tank is going to hit them in the mouth.
| Competitor | Public Pricing | Actual Pricing (Intel) | Pricing Model | Sales Motion |
|---|---|---|---|---|
| GovEagle | None published | Unpublished (enterprise) | Enterprise custom | Demo required, 1-week onboarding |
| GovDash | None published | Custom, modular | Seat-free, module-based | Demo required |
| CLEATUS | Published | $39/mo (DATA), $78/mo (DATA+AI) | Self-service | Walk-up |
| GovSignals | None published | Custom enterprise | Demo required | Sales-led |
| Vultron | None published | Custom (raised $22M Series A) | Demo required | Sales-led |
| Rohirrim | None published | Custom enterprise | Demo required | Sales-led |
| GovWin (Deltek) | None published | ~$40,000/year | Enterprise | Full sales cycle |
| GovTribe | None published | ~$25,000/year | Enterprise | Full sales cycle |
Observation: Only CLEATUS publishes walk-up pricing in the self-service segment. Every other player gates pricing behind a sales call. RFP Tank and its published tier structure will be immediately differentiated from every enterprise competitor the moment the pricing page goes live.
| Data Source | Cost | Data Available | API Quality |
|---|---|---|---|
| SAM.gov | Free (API key required) | All federal solicitations, contract awards, entity registrations | Well-documented, REST, reliable |
| USASpending.gov | Free | Award data, historical pricing, agency spend by NAICS | Good, bulk download available |
| FPDS (Federal Procurement Data System) | Free | Contract awards, modifications | Older API, some quirks |
| State procurement portals | Free (scraping) | Variable by state | No standard API; scraping required for SLED coverage |
| Federal Register | Free | Regulatory notices, advance procurement notices | XML feeds available |
Total data infrastructure cost: $0. The public data is free. The value is not in the data - it is in what you do with it.
Proposal prepared by IT Pro Partner | Germaine Brown | August 2026
RFP Tank is an IT Pro Partner product. rfptank.com is confirmed available as of proposal date.
All market data sourced from SAM.gov, USASpending.gov, SBA.gov, and publicly available competitor research.
Captured live during ideation with Germaine. One entry per numbered
requirement as they are stated. This file is the source of truth for the
subscriber-facing behavior; the proposal (rfptank-proposal.md) and architecture
(v3.6-architecture.md) reference it.
A logged-in subscriber can submit their proposal for compliance review by either:
different code paths that must both converge on extracted, clean text before
compliance scoring runs. File upload = parse PDF/DOCX/etc. URL = fetch +
extract (PDF download, Google Docs/SharePoint export, or plain page).
already specs the RFP side ("Paste a SAM.gov URL, upload a PDF, or drop in a
Word document"). Requirement 1 is the PROPOSAL side. RFP Tank scores the
proposal against the RFP, so a full run needs BOTH ingested.
attachments, past performance. Ingestion must accept 1..N documents and treat
them as one submission, not a single file.
portals behind login) and non-extractable pages are the common failure. The
URL flow needs a clear "can't reach this link" error with a fallback to
manual upload, not a silent empty ingest.
step of the journey, so auth (shared Stack Auth/Clerk tenant) has to exist
before any ingestion is usable. This is a dependency on the shared auth layer
already decided for ProposalTank + VerdictTank + RFP Tank.
Two parts, both from Germaine's second requirement message.
How many NAICS codes a subscriber can attach to their profile. Purpose: NAICS
is the primary filter for RFP matching/tracking (which opportunities are
relevant to this subscriber).
~99% under 10. Keeps matching signal clean.
unlimited.
business description and auto-suggest related codes from the NAICS hierarchy.
When a subscriber is tracking an RFP and that RFP is amended or changed, send
them a notification. Table-stakes feature (all competitors do amendment
alerts; a missed amendment = late or non-compliant proposal).
SAM.gov Opportunities API for that notice on a schedule (every 6-12 hrs),
diffs against last-seen version, notifies on change.
set-aside changed, requirements/description updated.
as a premium add.
requirement), not just "an amendment was posted."
Frame: the gold standard is not a longer feature list, it is depth on ONE
thing competitors do shallowly - traceability - plus an evaluation layer they
do not do at all. RFP Tank owns everything UP TO writing; ProposalTank writes
the prose; VerdictTank reviews the draft. No feature bleed.
mapped to the exact proposal paragraph that answers it, with confidence and
quoted evidence. Exportable to XLSX/PDF. This is the deliverable a pursuit
manager prints and hands to leadership.
thin coverage = weakness; separate the 3 real risks from the 11 fine items.
wording (evaluators keyword-scan) and where it is vague.
RFP's own evaluation criteria and weights. Answers "will it win", not just
"does it comply". Uses the inherited VerdictTank panel engine.
addressed, each with evidence.
signatures present, naming conventions. Automated gate. Top reason
proposals are rejected unread.
each requirement pre-mapped to a section and a page budget per section.
Writer fills in, never starts blank.
updates as they type.
requirements and answers are now stale (extends Req 2b).
| Dimension | Score | Type |
|---|---|---|
| Problem Definition | 7 | idea |
| Market Analysis | 3 | idea |
| Competitive Moat | 3 | idea |
| Business Model | 4 | proposal |
| Unit Economics | 3 | proposal |
| Technical Architecture | 4 | proposal |
| Go-to-Market Strategy | 5 | proposal |
| Risk Assessment | 6 | proposal |
| Execution Feasibility | 3 | proposal |
| Growth Trajectory | 2 | proposal |
Status: Build specification, companion to the v3.6 business proposal.
Scope: Extends the v3 architecture (5-stage pipeline, 4 API endpoints, 10 data model tables) with the 9 new v3.6 capabilities: Chat-to-Refine, URL-to-Review, Dual Scoring, Explain-the-Low-Score, Fix-It Remediation, Second-Opinion Audit Agent, Configurable Rules Engine, White-Label for Consultants, and Roast/Boost Community Peer Review.
Audience: Engineering. This is the build spec, not the pitch.
v3.6 scope note (Priority Fix #6 / Judge 3 Week 1 Cut List): The critical review of v3.6 found the 15-service, 21-table build scoped "like a multi-engineer build" for a solo-founder timeline. This document now describes two tiers: the 8-Week MVP (5 services, shipped) and Deferred (Post-Launch) components (5 services, designed but not built until a funnel/demand signal justifies them). Deferred components' schemas and designs are preserved below for continuity - cutting scope does not mean deleting the design work, it means sequencing it after MVP validation. See Section 9 for the revised build sequence.
βββββββββββββββββββββββββββββββββββββββββββββββ
β CLIENTS β
β Web App Β· Mobile Β· API Consumers β
βββββββββββββββββββββ¬ββββββββββββββββββββββββββββ
β HTTPS
ββββββββββββββββββββββΌβββββββββββββββββββββββββ
β EDGE / API GATEWAY (Caddy) β
β TLS term Β· rate limit Β· auth (JWT/API key) β
β request logging β
βββββββββ¬ββββββββββββββββββββββββββββββββββββββββ
β
βββββββββββββββββββββββΌββββββββββββββββββββββββ
β APPLICATION SERVICES (MVP: 2 of 5 kept) β
β review-api Β· chat-api β
βββββββββββββ¬ββββββββββββββββββββββββββββββββββββ
β enqueue
βββββββββββββΌββββββββββββββββββββββββββββββββββββββββββββββββ
β JOB QUEUE (Redis + RQ/Celery) β
β review.pipeline Β· chat.turn Β· url.extract β
β remediation.generate Β· sanitize.scan β
βββββββββββββ¬ββββββββββββββββββββββββββββββββββββββββββββββββββ
β
βββββββββββββββββββββββββββββΌβββββββββββββββββββββββββββββββββββββββββββββββββββ
β REVIEW PIPELINE WORKERS (pipeline-worker) β
β β
β Phase 1 Phase 2 Phase 3 Phase 4 Phase 5 β
β Research βββΆ Primary βββΆ Validation βββΆ Cross-Check βββΆ Verdict β
β Agent Reviewer Reviewer A + B (β) Aggreg. β
β (web verify (10-dim, (challenges (independent (majority β
β + citations) 1-10 + tags primary score) re-score) + dual β
β idea/proposal) scores) β
β β
β [Phase 6 - Audit Agent: DEFERRED, see Β§1.2 and Β§9 Post-Launch Track] β
βββββββββββββ¬ββββββββββββββββββββββββββββββββββββββββββββββ¬βββββββββββββββββββββββ
β β
βββββββββββββΌββββββββββββββββ ββββββββββββββΌβββββββββββββββββββ
β SANITIZATION GATE β β POST-PIPELINE GENERATORS β
β scans: verdict JSON, β β Explainer (per-dim text) β
β explanations, remediation,β β Fix-It Generator (action plan)β
β chat transcripts β β (Rules Engine overlay: DEFERRED)β
β before egress β ββββββββββββββββ¬ββββββββββββββββββββ
βββββββββββββ¬ββββββββββββββββ β
β β
βββββββββββββΌβββββββββββββββββββββββββββββββββββββββββββββββββΌββββββββββββββββ
β DATA LAYER β
β PostgreSQL (primary OLTP, org_id column present but multi-tenant enforcement β
β (RLS) deferred until White-Label track - see Β§1.2) β
β pgvector (corpus embeddings/similarity) β
β Redis (queue, session cache, rate limits, feature flags) β
β S3-compatible object store (PDFs, chat exports, uploaded docs, extracted HTML)β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β SUPPORTING SERVICES (MVP) β
β url-extractor (Crawl4AI) Β· fixit-generator Β· pdf-generator β
β cron (prediction T+90/180/365) β
β [domain-verifier, moderation-queue: DEFERRED - see Β§1.2] β
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
Per Priority Fix #6, the 8-week solo-founder MVP ships 5 services; the remaining 5 services from the original v3.6 scope are deferred to post-launch and gated on real usage/demand signal, not built speculatively. This directly implements the review's Week 1 Cut List (KEEP: Dual Scoring, Explain, Fix-It MVP, URL-to-Review, Corpus Schema / CUT: Audit Agent, Rules Engine, White-Label, Roast/Boost, and Chat-to-Refine deferred-not-deleted per below).
MVP - Ship in 8 Weeks (5 services)
| Layer | Component | MVP Status | Notes |
|---|---|---|---|
| App | review-api |
KEEP | Core review submission/status/verdict endpoints; dual-score fields, /review/url |
| App | chat-api |
KEEP (reduced) | Chat-to-Refine, but scoped down to the ghostwriting-guard behavior in Β§5.1 - no chat-ws streaming socket in MVP, REST polling only, to cut infra surface. Positioned as the "un-copilot" (see Β§5.1) rather than deferred outright, since Judge 3 flagged it as a differentiator once guarded correctly - see rationale note below. |
| Worker | url-extractor |
KEEP | Crawl4AI-based content extraction + sanitize; Feature 2 (URL-to-Review) |
| Worker | pipeline-worker |
KEEP | Phases 1-5 (Research β Primary β Validation β Cross-Check β Verdict/Dual-Score Aggregation). Phase 6 (Audit Agent) removed from the worker in MVP. |
| Worker | fixit-generator |
KEEP | Fix-It Remediation, upgraded to quality-scoring per Β§5.5 - retention-loop feature, explicitly called out as a KEEP in the review's cut list |
Note on Chat-to-Refine placement: The review's Week 1 Cut List names Chat-to-Refine as a CUT ("funnel feature; build when you have a funnel"). This document keeps a minimalchat-apiin MVP scope specifically because the ghostwriting guard in Β§5.1 is cheap to build (a length/format check, not new infrastructure) and de-risks the single biggest brand-reputation exposure identified by Judge 3 if Chat-to-Refine ships at all, in this MVP or later. If engineering capacity is tighter than modeled,chat-apiis the first MVP component to cut - in which case Chat-to-Refine moves to the Deferred table below in its entirety,chat-wsand all real-time streaming remain deferred either way, and the guard design in Β§5.1 is preserved for whenever the feature ships.
DEFERRED - Post-Launch, Built on Demand Signal (5 services)
| Layer | Component | MVP Status | Notes | Deferred Trigger |
|---|---|---|---|---|
| App | rules-api |
CUT (deferred) | CRUD + validation for org rule definitions | Revisit once an Enterprise/White-Label deal is in active negotiation and requires it |
| App | org-api |
CUT (deferred) | white-label org config, branding, API keys | Blocked on Minimum Viable Legal (MSA + Privacy Program) per the review's Legal Killers - see Β§4 |
| App | peer-review-api |
CUT (deferred) | Roast/Boost submission + moderation | Moderation overhead + defamation exposure (Judge 4 Legal Killer #4) outweighs launch value |
| Worker | audit-agent-worker |
CUT (deferred) | Phase 6 of pipeline, second-opinion audit | "Enterprise trust problem, not launch problem" per Judge 3; revisit once Enterprise tier has real pilot demand |
| Worker | moderation-worker |
CUT (deferred) | abuse handling for Roast/Boost | Dependent on peer-review-api; deferred together |
| Infra | domain-verifier |
CUT (deferred) | CNAME/TXT verification for white-label custom domains | Dependent on org-api / White-Label track; also blocked on Minimum Viable Legal |
The rules-evaluator overlay logic (Β§5.7) and Rules Engine pipeline injection point are deferred along with rules-api - no schema or prompt-injection surface for org rules ships in MVP. The full designs for all deferred components remain documented in Sections 5.6-5.9 and 7 of this document as the intended post-launch build, not as speculative dead weight - they represent validated architecture pending a demand signal, not throwaway work.
MVP Data Layer implications: Row-level security (RLS) multi-tenant policies (Β§6.3, Β§7.4), the org_configs/org_rules/org_api_keys/rule_templates tables, and cross-org isolation logic are not required for MVP since org-api and rules-api are deferred - org_id columns remain present (nullable) on core tables for forward compatibility, but RLS enforcement, domain routing, and multi-tenant billing are built when the White-Label track is actually resumed, gated on Minimum Viable Legal per Section 4.
users ββ1:Nββ reviews ββ1:1ββ verdicts ββ1:Nββ dimensions ββ1:1ββ dimension_explanations β β β β β ββ1:Nββ remediation_plans (per dimension) β βββ1:Nββ judge_scores ββN:1ββ judges β βββ1:Nββ research_citations β βββ1:Nββ predictions β βββ1:Nββ audit_findings (v3.6) β βββ1:Nββ peer_reviews (v3.6) β βββ1:1ββ chat_sessions (optional, pre-review link) (v3.6) β βββ1:Nββ chat_sessions ββ1:Nββ chat_messages (v3.6) β βββN:1ββ org_configs ββ1:Nββ org_rules ββN:1ββ rule_templates (v3.6, enterprise/white-label) corpus_entries (derived from verdicts, searchable, org_id-scoped) audit_log (cross-cutting, all tables)
All new v3.6 tables carry org_id (nullable for individual/non-enterprise accounts) to support multi-tenant partitioning for White-Label. All v3 tables are extended with org_id as part of the Phase 0 migration described in Section 9.
Dual-Score Schema Validation Gate (Priority Fix #2 / Judge 3 & Judge 1 finding, Architecture Β§2.2 lines 147-148 in the v3.6 review): The original v3.6 design bakedidea_score/proposal_scoredirectly into thereviewsandcorpus_entriestables as permanent NUMERIC columns from day one - a schema bet made before any beta user confirmed that a two-score split is more useful than the v3 single composite score. The review's Priority Fix #2 is explicit: "Do not bakeidea_score/proposal_scoreinto the corpus before validating with beta users that two scores are wanted. Gate the migration behind user testing, or ship dual-scoring as a non-schema overlay first."
>
v3.6 resolution - ship as overlay, migrate only after validation:
1. MVP ships dual scoring as a JSONB overlay, not dedicated columns. Theidea_score/proposal_scorevalues computed at Phase 5 (see Β§5.3) are written intoverdicts.verdict_json(already a JSONB column, no migration required) under adual_scorekey, and into a newreviews.dual_score_overlay JSONBcolumn (nullable, additive, zero-downtime to add/drop) rather than as first-class typedNUMERIC(5,2)columns onreviewsorcorpus_entries.
2. The NUMERIC columns shown below (idea_score,proposal_scoreonreviews; same oncorpus_entries, Β§2.3) are the target schema for GA, not the MVP schema. They are documented here for architectural continuity but are markedDEFERRED - VALIDATION GATEand must not be created by the Phase 0 migration until the exit criterion below is met.
3. Validation exit criterion (must pass before the permanent-column migration runs): a minimum of 25 beta users from the existing $29/mo Inner Circle cohort have used dual-score overlay output across at least 2 reviews each, and post-review survey/interview signal shows a majority preference for two scores over the legacy single composite. This reuses the same beta cohort the review's Priority Fix #1 calls for surveying - one survey instrument serves both purposes.
4. If validation fails (users prefer one score, or show no measurable preference), thedual_score_overlayJSONB column is dropped,verdict_json.dual_scorecontinues to carry the value for backward-compatible display only, and the corpus indexing strategy in Β§2.3/Β§2.4 reverts tolegacy_composite_scoreas the primary sortable/filterable column.
5. Corpus search and percentile lookups (Β§6.4) run against the JSONB overlay during the validation window - idx_corpus_dual_score_overlay is a GIN index on the JSONB path, functionally equivalent to a B-tree NUMERIC index for the query patterns in Β§3.2/Β§3.3 but avoids a schema commitment. Query latency is slightly higher (GIN vs. B-tree) but acceptable at MVP corpus volume; this is re-benchmarked before the permanent migration.
-- USERS (v3, extended)
CREATE TABLE users (
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
org_id UUID REFERENCES org_configs(id), -- v3.6: multi-tenant scoping
email TEXT UNIQUE NOT NULL,
password_hash TEXT,
tier TEXT NOT NULL DEFAULT 'free', -- free|pro|enterprise|white_label|beta
role TEXT NOT NULL DEFAULT 'member', -- member|org_admin|consultant|superadmin
created_at TIMESTAMPTZ NOT NULL DEFAULT now(),
updated_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
CREATE INDEX idx_users_org ON users(org_id);
-- REVIEWS (v3, extended: dual score columns, submission source, org scoping)
CREATE TABLE reviews (
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
org_id UUID REFERENCES org_configs(id),
user_id UUID NOT NULL REFERENCES users(id),
source_type TEXT NOT NULL DEFAULT 'file', -- v3.6: file|url|api
source_url TEXT, -- v3.6: populated when source_type='url'
chat_session_id UUID REFERENCES chat_sessions(id), -- v3.6: link to pre-review coaching
vertical TEXT,
status TEXT NOT NULL DEFAULT 'queued', -- queued|running|audit|done|failed
idea_score NUMERIC(5,2), -- DEFERRED - VALIDATION GATE: do not populate until dual-score validation exit criterion (Β§2.2) is met; MVP writes to dual_score_overlay instead
proposal_score NUMERIC(5,2), -- DEFERRED - VALIDATION GATE: see idea_score note above
dual_score_overlay JSONB, -- v3.6 MVP: { "idea_score": 78.5, "proposal_score": 41.0 }, non-schema overlay per Β§2.2 gate; promoted to typed columns above only post-validation
legacy_composite_score NUMERIC(5,2), -- v3: preserved for backward compat
rules_applied JSONB DEFAULT '[]', -- DEFERRED - org_rules snapshot, populated only once Rules Engine (Β§5.7) ships post-launch
parent_review_id UUID REFERENCES reviews(id), -- v3.6 (Β§5.5): set when this review is a resubmission addressing a prior remediation_plan
fix_completion_score NUMERIC(3,2), -- v3.6 (Β§5.5): mean fix_quality_tier (0-3) across parent's action_items, computed on resubmission
created_at TIMESTAMPTZ NOT NULL DEFAULT now(),
completed_at TIMESTAMPTZ
);
CREATE INDEX idx_reviews_org ON reviews(org_id);
CREATE INDEX idx_reviews_user ON reviews(user_id);
CREATE INDEX idx_reviews_status ON reviews(status);
CREATE INDEX idx_reviews_parent ON reviews(parent_review_id);
-- VERDICTS (v3, unchanged shape, verdict_json now carries dual-score payload)
CREATE TABLE verdicts (
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
review_id UUID NOT NULL REFERENCES reviews(id) UNIQUE,
verdict_json JSONB NOT NULL, -- see Section 2.5 for v3.6 schema
pdf_url TEXT,
public_share_id TEXT UNIQUE,
created_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
-- DIMENSIONS (v3, extended: score_type tag for dual scoring)
CREATE TABLE dimensions (
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
review_id UUID NOT NULL REFERENCES reviews(id),
name TEXT NOT NULL, -- e.g. "Market Analysis"
score NUMERIC(4,2) NOT NULL, -- 1-10, existing v3 scale
score_type TEXT NOT NULL DEFAULT 'proposal', -- v3.6: 'idea' | 'proposal'
weight NUMERIC(4,3) DEFAULT 1.0,
created_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
CREATE INDEX idx_dimensions_review ON dimensions(review_id);
-- JUDGES (v3, unchanged)
CREATE TABLE judges (
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
role TEXT NOT NULL, -- Research|Primary|Validation|CrossCheckA|CrossCheckB|...
vendor_internal TEXT NOT NULL, -- stripped by sanitization gate before egress
is_audit_agent BOOLEAN NOT NULL DEFAULT false, -- v3.6: flags the Phase 6 role
active BOOLEAN NOT NULL DEFAULT true,
accuracy_score NUMERIC(5,4),
created_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
-- JUDGE_SCORES (v3, unchanged)
CREATE TABLE judge_scores (
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
review_id UUID NOT NULL REFERENCES reviews(id),
judge_id UUID NOT NULL REFERENCES judges(id),
dimension_id UUID REFERENCES dimensions(id),
raw_score NUMERIC(4,2) NOT NULL,
rationale TEXT,
created_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
-- RESEARCH_CITATIONS (v3, unchanged)
CREATE TABLE research_citations (
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
review_id UUID NOT NULL REFERENCES reviews(id),
source_url TEXT NOT NULL,
excerpt TEXT,
retrieved_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
-- PREDICTIONS (v3, unchanged)
CREATE TABLE predictions (
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
review_id UUID NOT NULL REFERENCES reviews(id),
predicted_risk TEXT NOT NULL,
check_at TIMESTAMPTZ NOT NULL, -- T+90/180/365
outcome TEXT, -- pending|materialized|avoided
checked_at TIMESTAMPTZ
);
-- CORPUS_ENTRIES (v3, extended: org scoping for white-label segmentation, dual score overlay)
CREATE TABLE corpus_entries (
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
org_id UUID REFERENCES org_configs(id), -- v3.6: NULL = shared/general corpus
review_id UUID NOT NULL REFERENCES reviews(id),
vertical TEXT,
idea_score NUMERIC(5,2), -- DEFERRED - VALIDATION GATE, see Β§2.2; not populated at MVP
proposal_score NUMERIC(5,2), -- DEFERRED - VALIDATION GATE, see Β§2.2; not populated at MVP
dual_score_overlay JSONB, -- v3.6 MVP: populated instead of the two columns above, see Β§2.2
embedding VECTOR(1536),
is_public BOOLEAN NOT NULL DEFAULT false,
created_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
CREATE INDEX idx_corpus_org ON corpus_entries(org_id);
CREATE INDEX idx_corpus_embedding ON corpus_entries USING hnsw (embedding vector_cosine_ops);
CREATE INDEX idx_corpus_dual_score_overlay ON corpus_entries USING gin (dual_score_overlay); -- v3.6 MVP: GIN index serves overlay queries until/unless promoted to typed columns, see Β§2.2
-- AUDIT_LOG (v3, extended: new event types for v3.6 surfaces)
CREATE TABLE audit_log (
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
org_id UUID REFERENCES org_configs(id),
actor_id UUID REFERENCES users(id),
event_type TEXT NOT NULL, -- + v3.6: chat.message, review.url_submit, rule.create,
-- audit_agent.finding, peer_review.submit, org.brand_update
entity_type TEXT,
entity_id UUID,
metadata JSONB DEFAULT '{}',
created_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
CREATE INDEX idx_audit_log_entity ON audit_log(entity_type, entity_id);
-- CHAT_SESSIONS (Feature 1: Chat-to-Refine)
CREATE TABLE chat_sessions (
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
user_id UUID NOT NULL REFERENCES users(id),
org_id UUID REFERENCES org_configs(id),
status TEXT NOT NULL DEFAULT 'active', -- active|archived|converted_to_review
review_id UUID REFERENCES reviews(id), -- set once user submits for review
title TEXT,
turn_count INT NOT NULL DEFAULT 0,
created_at TIMESTAMPTZ NOT NULL DEFAULT now(),
updated_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
CREATE INDEX idx_chat_sessions_user ON chat_sessions(user_id);
-- CHAT_MESSAGES
CREATE TABLE chat_messages (
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
session_id UUID NOT NULL REFERENCES chat_sessions(id),
role TEXT NOT NULL, -- user|coach
content TEXT NOT NULL,
sanitized BOOLEAN NOT NULL DEFAULT false, -- gate must run before export
created_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
CREATE INDEX idx_chat_messages_session ON chat_messages(session_id, created_at);
-- DIMENSION_EXPLANATIONS (Feature 4: Explain the Low Score)
CREATE TABLE dimension_explanations (
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
dimension_id UUID NOT NULL REFERENCES dimensions(id) UNIQUE,
explanation_text TEXT NOT NULL,
cited_gaps JSONB NOT NULL DEFAULT '[]', -- ["no TAM calculation", "no competitor pricing"]
sanitized BOOLEAN NOT NULL DEFAULT false,
generated_by UUID REFERENCES judges(id), -- always the Primary Reviewer judge
created_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
-- REMEDIATION_PLANS (Feature 5: Here's How to Fix It - quality-scored per Β§5.5)
CREATE TABLE remediation_plans (
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
review_id UUID NOT NULL REFERENCES reviews(id),
dimension_id UUID NOT NULL REFERENCES dimensions(id),
action_items JSONB NOT NULL DEFAULT '[]',
-- action_items shape: [{ "description": str, "difficulty": 1-3,
-- "estimated_time_minutes": int, "template_url": str|null,
-- "fix_quality_tier": int|null, -- 0=not_attempted,1=superficial,2=minimal,3=substantive; NULL until resubmission evaluated, see Β§5.5
-- "fix_quality_rationale": str|null, -- evaluator's one-sentence justification, sanitized per Β§7.2
-- "evaluated_at": timestamptz|null }]
tier_scope TEXT NOT NULL DEFAULT 'full', -- 'summary' (Free) | 'full' (Pro+)
disclaimer_version TEXT NOT NULL DEFAULT 'DISC-001', -- v3.6: legal disclaimer version active at generation time, see Β§2.6
sanitized BOOLEAN NOT NULL DEFAULT false,
created_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
CREATE INDEX idx_remediation_review ON remediation_plans(review_id);
-- Fix quality aggregate lives on REVIEWS (added to the reviews table above, Β§2.2):
-- reviews.fix_completion_score NUMERIC(3,2) -- mean fix_quality_tier (0-3) across all action_items on resubmission, see Β§5.5
-- reviews.parent_review_id UUID REFERENCES reviews(id) -- links a resubmission to the review whose remediation plan it's addressing
-- AUDIT_FINDINGS (Feature 6: Second-Opinion Audit Agent) [DEFERRED - Post-Launch, see Β§1.2/Β§5.6]
-- Table shape preserved here for continuity; not created by the MVP Phase 0 migration.
CREATE TABLE audit_findings (
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
review_id UUID NOT NULL REFERENCES reviews(id),
finding_type TEXT NOT NULL, -- blind_spot|groupthink|underweighted_dim|tagging_inconsistency|no_finding
description TEXT NOT NULL,
severity TEXT NOT NULL DEFAULT 'low', -- low|medium|high
related_dimension_id UUID REFERENCES dimensions(id),
disclaimer_version TEXT NOT NULL DEFAULT 'DISC-001', -- v3.6: see Β§2.6
sanitized BOOLEAN NOT NULL DEFAULT false,
created_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
CREATE INDEX idx_audit_findings_review ON audit_findings(review_id);
-- RULE_TEMPLATES (Feature 7: Configurable Review Rules Engine) [DEFERRED - Post-Launch, see Β§1.2/Β§5.7]
-- Table shape preserved here for continuity; not created by the MVP Phase 0 migration.
CREATE TABLE rule_templates (
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
name TEXT NOT NULL,
category TEXT NOT NULL, -- compliance|brand_voice|required_field|checklist
schema_json JSONB NOT NULL, -- JSON Schema describing valid rule params
is_system BOOLEAN NOT NULL DEFAULT true, -- system-provided vs org-authored
created_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
-- ORG_RULES [DEFERRED - Post-Launch, see Β§1.2/Β§5.7]
CREATE TABLE org_rules (
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
org_id UUID NOT NULL REFERENCES org_configs(id),
template_id UUID REFERENCES rule_templates(id),
name TEXT NOT NULL,
rule_definition JSONB NOT NULL,
-- rule_definition shape: { "condition": {...if/then AST...},
-- "required_fields": [...], "checklist": [...], "severity": "block|warn" }
is_active BOOLEAN NOT NULL DEFAULT true,
validated_at TIMESTAMPTZ, -- set once syntax + schema validation passes
created_by UUID REFERENCES users(id),
created_at TIMESTAMPTZ NOT NULL DEFAULT now(),
updated_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
CREATE INDEX idx_org_rules_org ON org_rules(org_id) WHERE is_active = true;
-- ORG_CONFIGS (Feature 8: White-Label for Consultants)
CREATE TABLE org_configs (
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
name TEXT NOT NULL,
tier TEXT NOT NULL DEFAULT 'white_label', -- enterprise|white_label
domain TEXT UNIQUE, -- custom domain, e.g. reviews.acceleratorx.com
domain_verified BOOLEAN NOT NULL DEFAULT false,
domain_verified_at TIMESTAMPTZ,
logo_url TEXT,
brand_colors JSONB DEFAULT '{}', -- { "primary": "#...", "accent": "#..." }
email_template TEXT, -- HTML template with token placeholders
corpus_segment TEXT NOT NULL DEFAULT 'shared', -- 'shared' | 'dedicated'
api_key_prefix TEXT UNIQUE,
created_at TIMESTAMPTZ NOT NULL DEFAULT now(),
updated_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
-- ORG_API_KEYS (supports Feature 8's per-org API key scoping)
CREATE TABLE org_api_keys (
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
org_id UUID NOT NULL REFERENCES org_configs(id),
key_hash TEXT NOT NULL UNIQUE,
label TEXT,
scopes JSONB DEFAULT '["review:write","review:read"]',
revoked_at TIMESTAMPTZ,
created_at TIMESTAMPTZ NOT NULL DEFAULT now(),
last_used_at TIMESTAMPTZ
);
-- PEER_REVIEWS (Feature 9: Roast/Boost Community Peer Review)
CREATE TABLE peer_reviews (
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
review_id UUID NOT NULL REFERENCES reviews(id),
user_id UUID NOT NULL REFERENCES users(id),
type TEXT NOT NULL, -- roast|boost
comment TEXT,
visibility TEXT NOT NULL DEFAULT 'submitter_only', -- submitter_only|shared
flagged BOOLEAN NOT NULL DEFAULT false,
flagged_reason TEXT,
moderation_status TEXT NOT NULL DEFAULT 'clean', -- clean|pending_review|removed
created_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
CREATE INDEX idx_peer_reviews_review ON peer_reviews(review_id);
CREATE UNIQUE INDEX idx_peer_reviews_dedupe ON peer_reviews(review_id, user_id);
-- PEER_REVIEW_RATE_LIMITS (supports 3/day/user cap)
CREATE TABLE peer_review_rate_limits (
user_id UUID PRIMARY KEY REFERENCES users(id),
review_date DATE NOT NULL,
count_today INT NOT NULL DEFAULT 0,
UNIQUE (user_id, review_date)
);
| # | Table | Origin | Purpose |
|---|---|---|---|
| 1 | users | v3 (extended) | account, tier, org membership |
| 2 | reviews | v3 (extended) | review lifecycle, dual scores, source type |
| 3 | verdicts | v3 | final verdict JSON, PDF, share link |
| 4 | dimensions | v3 (extended) | per-dimension score + idea/proposal tag |
| 5 | judges | v3 (extended) | panel roster incl. audit agent flag |
| 6 | judge_scores | v3 | raw per-judge scoring |
| 7 | research_citations | v3 | Phase 1 web verification sources |
| 8 | predictions | v3 | T+90/180/365 outcome tracking |
| 9 | corpus_entries | v3 (extended) | searchable review corpus, org-segmented |
| 10 | audit_log | v3 (extended) | cross-cutting event log |
| 11 | chat_sessions | v3.6 | Chat-to-Refine sessions |
| 12 | chat_messages | v3.6 | Chat-to-Refine turn history |
| 13 | dimension_explanations | v3.6 | Explain-the-Low-Score text |
| 14 | remediation_plans | v3.6 | Fix-It action plans |
| 15 | audit_findings | v3.6 | Second-Opinion Audit Agent output |
| 16 | rule_templates | v3.6 | reusable rule schemas |
| 17 | org_rules | v3.6 | org-specific custom rules |
| 18 | org_configs | v3.6 | white-label branding/domain |
| 19 | org_api_keys | v3.6 | per-org scoped API keys |
| 20 | peer_reviews | v3.6 | Roast/Boost entries |
| 21 | peer_review_rate_limits | v3.6 | abuse/rate control |
{
"review_id": "uuid",
"status": "done",
"idea_score": 78.5,
"proposal_score": 41.0,
"dual_score_overlay": { "idea_score": 78.5, "proposal_score": 41.0 },
"framing_summary": "Your idea is strong; your pitch needs work.",
"disclaimer": "This is an automated critique generated by AI models, not professional advice. Scores and commentary are opinions, not facts. Do not rely on this output for investment, procurement, legal, or funding decisions without independent professional review. See full disclaimer at https://verdicttank.com/legal/disclaimer.",
"dimensions": [
{
"name": "Market Analysis",
"score": 4,
"score_type": "idea",
"explanation": {
"text": "Market Analysis: 4/10 - no TAM calculation, no competitor pricing data, assumes zero competition.",
"cited_gaps": ["no TAM calculation", "no competitor pricing data", "zero-competition assumption"],
"disclaimer": "AI-generated critique, not professional advice."
},
"remediation": {
"tier_scope": "full",
"disclaimer": "This remediation plan is an automated suggestion, not professional or legal advice. Validate independently before acting.",
"quality_tier": "not_yet_evaluated",
"action_items": [
{
"description": "Calculate TAM using the top-down/bottom-up hybrid formula",
"difficulty": 2,
"estimated_time_minutes": 90,
"template_url": "https://cdn.verdicttank.com/templates/tam-calc.xlsx"
},
{
"description": "Add a competitor pricing comparison table",
"difficulty": 1,
"estimated_time_minutes": 45,
"template_url": "https://cdn.verdicttank.com/templates/competitor-pricing.xlsx"
}
]
}
}
],
"audit": {
"status": "deferred_post_mvp",
"findings": [],
"tagging_consistency_check": "not_run",
"note": "Second-Opinion Audit Agent (Phase 6) is deferred post-MVP per Β§1.2. Field shape is preserved for forward compatibility; populated once the Audit Agent worker ships. When populated, audit.disclaimer carries the same 'no professional advice' language as other surfaces."
},
"rules_applied": [],
"corpus_percentile": { "idea_score": 80, "proposal_score": 30 },
"public_share_id": "vt-8x2k",
"generated_at": "2026-08-11T00:00:00Z"
}
rules_applied ships as an always-empty array in MVP since the Rules Engine (rules-api) is deferred per Β§1.2 - the field is preserved in the schema for forward compatibility rather than removed, avoiding a breaking payload change when the Rules Engine ships post-launch.
All freeform text fields (explanation.text, action_items[].description, audit.findings[].description) are sanitized before this object leaves the pipeline - see Section 7.2. All disclaimer fields are not sanitizable/removable content - they are fixed legal boilerplate injected by the response serializer after sanitization, never model-generated, and cannot be stripped by any per-review customization (org branding, White-Label templates, tier). See Β§2.6 for the full disclaimer policy.
Origin: Judge 4's Legal Killer #2 (AI Liability / Defamation / Tortious Interference - ranked EXTREME) and Priority Fix #9 of the v3.6 review: "Every verdict, explanation, audit finding, and remediation action item must carry conspicuous language... This is the first line of defense against Rank 2 liability." No disclaimer, limitation-of-liability language, or indemnity existed anywhere in v3.6 prior to this revision - Section 15/Risk Assessment covered only technical/product risk, not the liability exposure of an AI system whose scores can be blamed for a lost deal, declined funding round, or reputational harm.
Standard disclaimer text (canonical copy, referenced by ID DISC-001 so all surfaces stay in sync on future legal-copy revisions):
"This is an automated critique generated by AI models, not professional advice. Scores, explanations, audit findings, and remediation suggestions are AI-generated opinions, not verified facts or expert judgments. Do not rely on this output for investment, procurement, legal, regulatory, or funding decisions without independent professional review. VerdictTank and IT Pro Partner disclaim liability for decisions made in reliance on this output. Full terms: https://verdicttank.com/legal/disclaimer."
Coverage - every output surface carries it, no exceptions:
| Output Surface | Placement | Enforcement Mechanism |
|---|---|---|
Verdict JSON (verdict_json) |
Top-level disclaimer field + per-dimension explanation.disclaimer + per-remediation remediation.disclaimer (see Β§2.5) |
Injected by the API/PDF serializer post-sanitization; schema validation rejects a verdict payload missing the top-level field (fail-closed, same enforcement pattern as explanation.text in Β§5.4) |
| PDF reports | Persistent footer on every page + prominent banner directly below the score header on page 1 | Hardcoded in the PDF template (not model-generated, not org-brand-overridable - see White-Label note below) |
| Public share reports (web) | Sticky banner above the fold, not a footnote or dismissible tooltip | Rendered server-side in the share-page template, not client-injectable/removable |
API responses (/verdict/{id}, /verdict/{id}/pdf, /reviews/{id}/remediation, /reviews/{id}/audit) |
disclaimer field present on every response object that carries scores, explanations, findings, or remediation |
Contract-tested: API response schema tests fail CI if any scored/explained/remediated response type omits the field |
| Chat-to-Refine transcripts | Session-start system message ("I'm a coach, not an advisor - nothing here is professional advice") + persistent footer in chat UI | Rendered client-side from a fixed string, reinforces the "un-copilot" brand framing in Β§5.1 |
| Audit findings (when the Audit Agent ships post-MVP, Β§1.2) | Same disclaimer field pattern as explanations/remediation, applied at design time so no retrofit is needed when Phase 6 ships |
Schema shape reserved in Β§2.5 now; enforcement added to the audit-findings serializer at build time |
| Corpus/percentile displays | Inline caveat: "Percentile rankings are relative to other AI-scored submissions, not a market or investment benchmark" | Rendered alongside corpus_percentile wherever it's displayed |
Non-negotiable design constraints:
org_configs.email_template/brand_colors theming (Β§5.8) is explicitly scoped to cannot touch disclaimer copy, placement, or prominence - this is enforced at the template-rendering layer (disclaimer blocks are rendered from a separate, non-themeable partial, injected after the brand template resolves). Applies once White-Label ships post-MVP; documented now so it's not an afterthought when org-api is un-deferred.DISC-001 is versioned; if legal counsel revises the disclaimer text (e.g. after the Minimum Viable Legal review the critical review calls for - MSA + Privacy Program + Content Safety Stack), the version bump is a config change, not a code change, and old stored verdicts retain the disclaimer version active at generation time (verdict_json.disclaimer_version) for auditability.auth-api (unchanged from v3).Authorization: Bearer vt_live_... header. Keys are scoped per-org (org_api_keys) - v3.6 adds org scoping on top of v3's flat per-user API keys.org_configs.domain) are resolved to org_id at the edge before hitting any endpoint; all endpoints below implicitly filter/write with that org_id.POST /api/verdicttank/review
Upload-based review submission (file). Response now includes dual score fields once complete.
// Response (200, once status=done)
{
"review_id": "uuid",
"idea_score": 78.5,
"proposal_score": 41.0,
"status": "done"
}
GET /api/verdicttank/status/{id} - unchanged polling contract; status enum extended with audit (Phase 6 in progress).
GET /api/verdicttank/verdict/{id}/pdf - unchanged; PDF generator now renders dual-score header and per-dimension explanation/remediation sections.
GET /api/verdicttank/corpus/search - unchanged query contract; results now include idea_score and proposal_score columns instead of a single blended score, and are org_id-scoped for white-label callers.
POST /api/verdicttank/review/url (Feature 2: URL-to-Review)
// Request
{ "url": "https://example.com/pitch-deck", "vertical_hint": "fintech" }
// Response (202 Accepted)
{ "review_id": "uuid", "status": "queued", "source_type": "url" }
Validation: URL scheme allowlist (http/https), DNS resolution check, max content size 10MB, extraction timeout 30s (hard cap 45s), robots.txt respected, content sanitized before entering the pipeline (see Section 7.2).
POST /api/verdicttank/chat/start (Feature 1)
// Request: { "initial_message": "I'm building a marketplace for..." }
// Response: { "session_id": "uuid", "reply": "Tell me more about who pays on this marketplace." }
POST /api/verdicttank/chat/{session_id}/message
// Request: { "message": "Both sides pay a transaction fee." }
// Response: { "reply": "...", "turn_count": 4 }
Also available as WSS /ws/verdicttank/chat/{session_id} for real-time streaming token-by-token (see Section 5.1). REST polling variant is the fallback for clients that can't hold a socket.
GET /api/verdicttank/chat/{session_id}/history
Returns sanitized transcript (chat_messages where sanitized=true). Export triggers gate scan if not already sanitized.
POST /api/verdicttank/rules (Feature 7, Enterprise/White-Label only)
{ "name": "SOC2 disclosure required", "template_id": "uuid",
"rule_definition": { "condition": {"if": "vertical == 'fintech'"}, "required_fields": ["compliance.soc2_status"], "severity": "block" } }
Response 201 on pass, 422 with schema violation details on rule-syntax failure.
GET /api/verdicttank/rules / PATCH /api/verdicttank/rules/{id} / DELETE /api/verdicttank/rules/{id} - standard CRUD, org-scoped.
POST /api/verdicttank/org/branding (Feature 8, White-Label)
{ "logo_url": "...", "brand_colors": {"primary": "#0B1B33"}, "email_template": "<html>...</html>" }
POST /api/verdicttank/org/domain - initiates CNAME/TXT verification (see Section 5.6).
POST /api/verdicttank/reviews/{id}/peer-review (Feature 9)
{ "type": "roast", "comment": "Your CAC assumption ignores paid acquisition entirely." }
Rate-limited to 3/user/day (enforced via peer_review_rate_limits, Redis-cached counter). Returns 429 past cap.
POST /api/verdicttank/reviews/{id}/peer-review/{peer_review_id}/report - abuse reporting, pushes into moderation-queue.
GET /api/verdicttank/reviews/{id}/remediation - standalone fetch of the Fix-It plan (also embedded in verdict JSON).
GET /api/verdicttank/reviews/{id}/audit - standalone fetch of audit findings (Enterprise tier).
INTAKE EGRESS
β β
βΌ β
ββββββββββββββββ ββββββββββββββββ ββββββββββββββββ ββββββββββββββββββββ β
β File Upload β β URL Extract β β Chat-to- β β API submission β β
β (existing) β β (Crawl4AI) β β Refine draft β β (Enterprise) β β
ββββββββ¬ββββββββ ββββββββ¬ββββββββ ββββββββ¬ββββββββ ββββββββββ¬ββββββββββββ β
ββββββββββββββββββββ΄βββββββββββββββββββ΄βββββββββββββββββββββ β
β β
normalize to standard intake format β
β β
βΌ β
ββββββββββββββββββββββββββββββββββ β
β [Rules Engine pre-check: β DEFERRED - Post-Launch β
β DEFERRED, see Β§1.2] β (see Β§1.2, Β§5.7) β
ββββββββββββββββββ¬βββββββββββββββββ β
β β
ββββββββββββββββββββββββββββββββββΌββββββββββββββββββββββββββββββββββββββ β
β PHASE 1: Research Agent β β
β Live web verification + citation gathering β β
ββββββββββββββββββββββββββββββββββ¬ββββββββββββββββββββββββββββββββββββββ β
β β
ββββββββββββββββββββββββββββββββββΌββββββββββββββββββββββββββββββββββββββ β
β PHASE 2: Primary Reviewer β β
β 10-dim brutal critique, 1-10 per dim β β
β + score_type tag (idea|proposal) per dim, overlay-only [v3.6, Β§2.2] β β
β + per-dimension explanation generated inline [v3.6, Feature 4] β β
β [org rule dimensions: DEFERRED - see Β§1.2, Β§5.7] β β
ββββββββββββββββββββββββββββββββββ¬ββββββββββββββββββββββββββββββββββββββ β
β β
ββββββββββββββββββββββββββββββββββΌββββββββββββββββββββββββββββββββββββββ β
β PHASE 3: Validation Reviewer - challenges primary score β β
ββββββββββββββββββββββββββββββββββ¬ββββββββββββββββββββββββββββββββββββββ β
β β
βββββββββββββββββββββββββ΄ββββββββββββββββββββββββ β
βΌ βΌ β
ββββββββββββββββββββ ββββββββββββββββββββ β
β PHASE 4: β β PHASE 4: β β
β Cross-Check A β (parallel) β Cross-Check B β β
βββββββββββ¬βββββββββββ βββββββββββ¬βββββββββββ β
βββββββββββββββββββββββββ¬ββββββββββββββββββββββββββ β
βΌ β
ββββββββββββββββββββββββββββββββββββ β
β PHASE 5: Majority Verdict β β
β aggregates scores, builds β β
β idea_score + proposal_score β β
β as dual_score_overlay [v3.6, Β§2.2] β β
ββββββββββββββββββββ¬ββββββββββββββββββββ β
β β
ββββββββββββββββββββΌββββββββββββββββββββ β
β [PHASE 6: Audit Agent - DEFERRED, β β
β Post-Launch, see Β§1.2 / Β§5.6] β β
ββββββββββββββββββββ¬ββββββββββββββββββββ β
β β
ββββββββββββββββββββΌββββββββββββββββββββ β
β POST-PIPELINE: Fix-It Generator β β
β (Pro+; lightweight summary on Free) β β
β quality-scored, 3-tier rubric [v3.6, Β§5.5]β β
ββββββββββββββββββββ¬ββββββββββββββββββββ β
β β
ββββββββββββββββββββΌββββββββββββββββββββ β
β SANITIZATION GATE (extended) β β
β scans verdict JSON, explanations, β β
β remediation text β β
β (audit findings: N/A, Phase 6 deferred) β β
ββββββββββββββββββββ¬ββββββββββββββββββββ β
β β
ββββββββββββββββββββΌββββββββββββββββββββ β
β DISCLAIMER INJECTION (post-sanitization) ββββββββββββββββββββββββ
β "no professional advice" on every β β
β scored/explained/remediated surface [Β§2.6]β β
ββββββββββββββββββββ¬ββββββββββββββββββββ β
β β
ββββββββββββββββββββΌββββββββββββββββββββ β
β PDF Generation + Corpus Write ββββββββββββββββββββββββ
β (dual-score overlay, org-segmented) β
βββββββββββββββββββββββββββββββββββββββββ
chat_sessions.review_id links the two for provenance, but the coach model never writes to dimensions, judge_scores, or verdicts.source_url + snapshot key) for auditability - if the source page changes or disappears, the review remains reproducible.score_type tag set at Phase 2 generation time; Phase 5 aggregation buckets dimension scores by tag and computes two weighted composites (0-100 scale) instead of one.audit_findings rows and validates score_type tagging consistency across judges.422-style rejection before spending pipeline cost), and (2) injected as additional evaluation dimensions inside the Phase 2 Primary Reviewer prompt so soft/scored rules show up as extra dimensions rows tagged source=org_rule.Priority Fix #4 (Judge 3): The v3.6 review found the original prompt-level ghostwriting mitigation ("system prompt says don't ghostwrite, flag suspicious output for manual review") reactive and non-deterministic - a coaching model instructed not to ghostwrite can still ghostwrite a full paragraph, and "flag for manual spot-check" catches it after the user already received it. Judge 3's recommendation: build a deterministic, real-time output-length guard that rejects multi-paragraph output before it reaches the user, and lean into an explicit "un-copilot" brand position - VerdictTank coaches by asking better questions, it does not write your pitch for you, and it says so out loud.
chat-api (REST, MVP) - see Β§1.2 for the MVP note that chat-ws streaming is deferred; the guard below is described for the REST turn-response path and applies identically if/when streaming ships.Real-time output-length guard (deterministic, not prompt-level):
chat-api request path immediately after the coaching-model call returns and before the response is persisted to chat_messages or sent to the client.? (i.e., it isn't phrased as a question at all - a purely declarative multi-sentence "here's what your pitch should say" response fails this check even if it's under the length cap).#), numbered lists with more than 2 items, or bullet lists with more than 2 items.chat_messages.guard_rejected BOOLEAN, chat_messages.guard_rejection_reason TEXT, chat_messages.retry_count INT) for the quarterly manual audit in item 6 below, and to catch coaching-model/prompt regressions early (a spike in guard-trip rate on a model update is itself an alert-worthy signal, not just an audit finding).chat-api, not the web frontend, so it can't be bypassed by a different client hitting the same endpoint."Un-copilot" brand positioning (Judge 3's recommendation, product + UX + copy, not just engineering):
chat_sessions + chat_messages. Sessions are resumable - a user can leave and come back; turn_count and updated_at support a "recent sessions" list in the UI.reviews row with chat_session_id set.guard_rejected/guard_rejection_reason log (item 4 above) rather than relying on manual spot-check alone as the only detection layer. Sanitization gate runs on any transcript before it can be exported via GET /chat/{id}/history. A "no professional advice" disclaimer is shown at session start and persists in the chat UI footer - see Β§2.6.url-extractor worker, backed by Crawl4AI (headless-browser-based extraction with readability/structure heuristics).422 with a specific reason code rather than a generic failure, since users need to know whether to try again or upload a file instead.score_type: "idea"|"proposal".idea and Judge B tags it proposal, that's a flagged tagging_inconsistency finding, not a silent average.idea; clarity of argument, evidence quality, internal consistency β proposal. The mapping is a configurable lookup table (dimension_score_type_defaults), not hardcoded per-judge, so it can be tuned without a redeploy.explanation.text and explanation.cited_gaps[] as mandatory fields alongside every dimension score - the API schema rejects a dimension object missing them (fail-closed, not fail-open, per the v3.6 proposal's transparency mandate).dimension_explanations, 1:1 with dimensions.Priority Fix #7 (Fix-It Quality Loophole): The v3.6 review found the original remediation-tracking design ("re-review score deltas... detect superficial template-filling vs. genuine improvement") was aspirational but not actually specified - there was no evaluation step that distinguished a user who genuinely added a TAM calculation from a user who pasted a single sentence into the relevant section just to make the presence-check pass. Presence-checking ("did the field get filled in") is trivially gameable and undermines the entire retention-loop premise of Fix-It as a credible improvement signal, not a checkbox exercise.
remediation_plans rows, one per low-scoring dimension (threshold configurable, default: dimensions scoring β€5/10 get a plan; Pro+ can request plans on any dimension).tier_scope='summary', no action_items array, just a rolled-up text field reusing the same table with an empty array + a summary_text companion column); Pro+ gets the full structured action_items[] with difficulty, time estimate, and optional template links.3-Tier Fix Quality Rubric (evaluation, not presence-checking):
On resubmission (re-review of a previously scored document, or a follow-up chat/URL submission linked via reviews.parent_review_id), each action_item that was addressed is evaluated against the specific gap it named - not merely against whether the target section now contains text. Evaluation runs as an additional judge-model call scoped narrowly to comparing the before/after content for a given action_item, using the original cited_gaps and description as the rubric anchor.
| Tier | Score | Definition | Scoring Rule |
|---|---|---|---|
| Superficial | 1 | The relevant section changed, but the specific gap named in cited_gaps is still unaddressed - e.g. a TAM number was added but with no visible methodology, source, or calculation shown; or text was added that mentions the topic without supplying the missing data/evidence. |
Triggers when the evaluator can find no calculation, citation, data point, or structural change that actually closes the gap - text presence alone does not clear this tier. |
| Minimal | 2 | The gap is nominally addressed but shallow - e.g. a TAM figure is present with a one-line methodology note, but no bottom-up/top-down hybrid breakdown, no source citation, no sensitivity range. Passes a "did they try" bar but not a "would this survive investor scrutiny" bar. | Triggers when the evaluator finds a genuine attempt that directly responds to the cited gap, but missing at least one of: methodology transparency, supporting data/citation, or the specific technique the original action_item.description recommended. |
| Substantive | 3 | The gap is closed with the rigor the original action item called for - e.g. TAM calculated via the recommended hybrid formula, source-cited, with a stated methodology and range. The addition would plausibly change an informed reader's assessment of that dimension. | Triggers when the evaluator confirms the specific technique/data named in the action item is present and internally consistent with the rest of the submission (a number that contradicts other stated figures elsewhere in the document does not qualify, even if superficially "complete"). |
| Not Attempted | 0 | No detectable change to the relevant section between submissions. | Default when the diffed section is identical or near-identical to the original. |
action_items.fix_quality_tier INT (0-3, per the table above), action_items.fix_quality_rationale TEXT (the evaluator's one-sentence justification, sanitized like any other freeform field per Β§7.2), action_items.evaluated_at TIMESTAMPTZ.reviews.fix_completion_score is computed on resubmission as the mean tier score across all action_items that had a prior plan, giving a single 0-3 "did they actually improve it" number surfaced back to the user and usable as a retention/engagement metric internally - replacing the vague "score deltas" framing from the prior draft with a defined, auditable computation.fix_completion_score - a user can get a higher idea_score/proposal_score on resubmission (because the panel re-evaluated the actual content) and a low fix_completion_score if the improvement, while score-moving, didn't specifically target the gaps the Fix-It plan called out; the two numbers are reported separately so users see both "did your score go up" and "did you fix what we told you to fix."action_items[] carries the remediation.disclaimer field per Β§2.6 - Fix-It output is a suggestion, not professional or legal advice, and this is stated on every surface where a plan is rendered (API, PDF, web).This component is designed but not built for the 8-week MVP. Per the review's Week 1 Cut List, the Audit Agent addresses "an enterprise trust problem, not a launch problem" - it is retained here in full so the design isn't lost, and is revisited once Enterprise-tier pilot demand justifies the build (see Β§1.2 Deferred Trigger table).
judge_scores, all dimension score_type tags.audit_findings rows (blind_spot, groupthink, underweighted_dim, tagging_inconsistency, or no_finding if nothing surfaces - explicitly logged rather than omitted, because a "never finds anything" audit agent is itself a signal per the proposal).judges.is_audit_agent=true, judges.accuracy_score computed the same way).This component is designed but not built for the 8-week MVP. Per the review's Week 1 Cut List, the Rules Engine is a White-Label/Enterprise dependency with no MVP demand signal. Retained here as the intended post-launch build.
rules-api for CRUD, rules-evaluator worker embedded in the pipeline pre-check and Phase 2 prompt injection.rule_templates.schema_json (JSON Schema) at save time - malformed rules are rejected with a 422 before they ever reach a live review.severity: "block" run before Phase 1 - e.g. a required compliance field missing halts the review immediately, avoiding wasted pipeline spend.severity: "warn" are appended to the Primary Reviewer prompt as additional evaluation dimensions (dimensions.name prefixed, source='org_rule'), scored 1-10 alongside the standard rubric.This component is designed but not built for the 8-week MVP, and additionally gated on Minimum Viable Legal (MSA + Privacy Program) per the review's Legal Killers before any Enterprise/White-Label customer is onboarded - see Β§2.6 and Β§1.2.
org-api (branding/config CRUD), domain-verifier (background job polling DNS for CNAME/TXT confirmation).org_id. Row-level security policies in PostgreSQL enforce org_id scoping at the database layer as a second line of defense behind application-layer filtering (belt-and-suspenders against a missed WHERE org_id = ... in a future query).domain_configs.domain β system issues a CNAME target + TXT verification token β domain-verifier polls DNS every 5 minutes for up to 48 hours β on match, domain_verified=true, edge gateway (Caddy) is updated (via API-driven config reload, not manual) to route that domain's TLS + org_id resolution.logo_url, brand_colors (CSS custom properties injected at render time), and email_template (Jinja/Handlebars token substitution) are applied at the PDF generator and email-notification layer - no separate white-label codebase, one codebase with per-org theming.org_configs.corpus_segment controls whether an org's reviews land in the shared corpus (default) or a dedicated segment invisible to corpus-search queries from other orgs - required for the "no more than 3 accounts until segmentation is independently verified" launch condition.This component is designed but not built for the 8-week MVP. Per Judge 4's Legal Killer #4 (defamation/moderation exposure), the moderation overhead and liability surface outweigh launch-stage value - retained here as the intended post-launch build once moderation tooling and Minimum Viable Legal are in place.
peer-review-api + moderation-queue worker.reviews gets an implicit visibility flag via a companion review_settings.peer_review_enabled column, or reuse reviews.rules_applied-style JSONB settings blob - recommend a dedicated peer_review_enabled BOOLEAN DEFAULT false column on reviews).submitter_only - peer feedback lands in the submitter's own dashboard; nothing is public unless the submitter explicitly re-shares it.peer_review_rate_limits table double-enforce the 3/user/day cap (Redis for fast-path rejection, Postgres for durable audit trail).flagged/flagged_reason/moderation_status columns feed a moderation queue; repeated flags against a user trigger a cooldown (application-layer, not schema-level). Feature-flagged (feature_flags.roast_boost_enabled) per the proposal's Condition 7 - off by default for every tier including migrated beta accounts until one full quarter of opt-in data is reviewed.ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ β Edge tier (2x app3-class instances, HA pair) β β Caddy - TLS termination, org-domain routing, rate limiting β ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ β App tier (autoscaling pool, 3-6 instances) β β review-api Β· chat-api Β· rules-api Β· org-api Β· peer-review-api β β chat-ws (sticky sessions via Redis pub/sub for horizontal scale) β ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ β Worker tier (autoscaling pool, scales with queue depth) β β pipeline-workers (Phases 1-6) Β· url-extractor Β· fixit-generator β β sanitize-scanner Β· pdf-generator Β· domain-verifier Β· cron β ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ β Data tier β β PostgreSQL 16 (primary + read replica) w/ pgvector extension β β Redis 7 (queue + cache + rate limits + feature flags) β β S3-compatible object store (Wasabi/S3) for PDFs, transcripts, β β URL snapshots, uploaded source docs β ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
services:
edge-caddy: {image: caddy:2, ports: ["443:443"]}
review-api: {build: ./services/review-api}
chat-api: {build: ./services/chat-api}
chat-ws: {build: ./services/chat-ws}
rules-api: {build: ./services/rules-api}
org-api: {build: ./services/org-api}
peer-review-api: {build: ./services/peer-review-api}
pipeline-worker: {build: ./workers/pipeline, deploy: {replicas: 4}}
url-extractor: {build: ./workers/url-extractor} # wraps Crawl4AI
fixit-generator: {build: ./workers/fixit-generator}
audit-agent-worker: {build: ./workers/audit-agent}
sanitize-scanner: {build: ./workers/sanitize-scanner}
pdf-generator: {build: ./workers/pdf-generator}
domain-verifier: {build: ./workers/domain-verifier}
moderation-worker: {build: ./workers/moderation}
cron: {build: ./workers/cron} # T+90/180/365 predictions
postgres: {image: pgvector/pgvector:pg16, volumes: ["pgdata:/var/lib/postgresql/data"]}
redis: {image: redis:7-alpine}
HNSW index on corpus_entries.embedding for approximate nearest-neighbor similarity search.org_id-bearing tables as of the v3.6 migration - application connects with a role that has org_id set via SET app.current_org_id per-request, and RLS policies filter automatically. This is the second enforcement layer for white-label tenant isolation described in Section 5.8.roast_boost_enabled, audit_agent_billing_live, etc.) so risky features can be killed instantly without a deploy.idea_score/proposal_score distributions) are recomputed on a 15-minute cadence and cached in Redis rather than computed per-request - avoids a full aggregate scan on every verdict render.org_configs row) is cached at the edge layer keyed by resolved domain, invalidated on org.brand_update audit events.This section addresses the single most common enterprise buyer question: "where does this run?" VerdictTank is not an abstract cloud service - it runs on bare-metal infrastructure owned and operated by IT Pro Partner, with two deployment models to match the client's risk and isolation requirements.
The solution runs on IT Pro Partner's existing production infrastructure (ITPP-INFRA), the same hardware and backup pipeline that powers ITPP's own operations.
| Layer | Hardware | Details |
|---|---|---|
| Edge | netcup RS 4000 (app3) | Caddy TLS termination, org-domain routing, rate limiting |
| App + Worker | netcup RS 4000 (app2, app3) | Containerized services (Docker), autoscaling within the pool |
| Data | netcup RS 4000 (app2) | PostgreSQL 16 + pgvector, Redis 7 |
| Object Storage | Wasabi S3 (shared bucket) | PDFs, chat exports, URL snapshots, uploaded docs |
| Backup | ITPP backup pipeline | Daily 2-5 AM S3 sync, versioning ON, 90-day retention; database WAL archiving every 15 minutes via hermes-live-sync β s3://itpropartner-backups/verdicttank/ |
Best for: Pro tier, early Enterprise pilots, non-regulated clients. Cost-efficient - leverages existing capacity with no dedicated hardware overhead. All infrastructure is documented and auditable under ITPP's existing SOC 2-type operational controls.
The solution runs on dedicated netcup or Hetzner bare-metal instances provisioned specifically for the client, with a dedicated S3 bucket and isolated backup pipeline.
| Layer | Hardware | Details |
|---|---|---|
| Edge | Dedicated netcup RS 2000/4000 or Hetzner CPX | Isolated Caddy instance, client-specific TLS |
| App + Worker | Dedicated netcup RS 4000 or equivalent | No shared compute with other VerdictTank tenants |
| Data | Dedicated PostgreSQL + Redis on the same instance or separate, per client requirements | pgvector included |
| Object Storage | Dedicated Wasabi S3 bucket | No cross-tenant object storage; bucket owned by client or ITPP per contract |
| Backup | Dedicated backup pipeline | Independent S3 bucket with versioning ON; dedicated cron schedule; restore testing included quarterly |
Best for: Enterprise clients with regulatory requirements (HIPAA, ITAR, FedRAMP-adjacent), White-Label resellers who need full data isolation guarantees, or any client whose compliance framework requires dedicated infrastructure with no shared-tenancy risk. All infrastructure is managed by IT Pro Partner - the client never touches servers - but the hardware, storage, and backup pipeline are theirs alone.
Shared responsibility line: In both models, ITPP manages everything below the application layer (OS, Docker, database, backups, monitoring, patching, incident response). The client's only responsibility is user account management and content submitted to the platform. This is the same operational model ITPP uses for all managed infrastructure - the client gets the benefit of bare-metal performance without any of the operational burden.
org_api_keys) for Enterprise/White-Label programmatic access - a compromised key can be scoped to review:read only, or revoked without affecting other keys on the same org.org_admin can manage org_rules, org_configs, and view portfolio-wide corpus; consultant role (White-Label) can manage client portfolios but not billing; member can only see their own reviews unless peer_review.visibility='shared'.The v3 gate blocked deploys unless a pre-deploy scan confirmed no internal vendor/model identity or architecture detail could leak into public reports, PDFs, or API responses. v3.6 quadruples the freeform-text surface area the gate must cover:
| Surface | v3 | v3.6 |
|---|---|---|
| Public share reports | β | β |
| API responses | β | β |
| PDF reports | β | β |
Dimension explanations (dimension_explanations.explanation_text) |
- | β new |
Remediation action items (remediation_plans.action_items[].description) |
- | β new |
Chat transcripts on export (chat_messages.content) |
- | β new |
Audit findings (audit_findings.description) |
- | β new |
sanitized=true and become eligible for egress. Fail-closed: unsanitized content is never served, regardless of tier or urgency.Two new v3.6 features introduce attacker-controlled or third-party text into a model context, which is new relative to v3's fully-controlled prompt pipeline:
rule_templates.schema_json, then rendered into a fixed prompt template slot (e.g. "Additionally verify: {field_name} is present") - the rule author cannot inject arbitrary instructions, only populate whitelisted template variables.org_id-bearing table is required (via a lint rule / query-builder wrapper, not convention alone) to include an org_id filter derived from the authenticated session/API key - never from a client-supplied parameter.corpus_entries.org_id + org_configs.corpus_segment='dedicated' keeps White-Label client data out of shared corpus-search results and out of cross-corpus benchmark comparisons unless the org explicitly opts into shared benchmarking.| Cost Component | v3 Baseline | v3.6 Delta | Running Total |
|---|---|---|---|
| Base pipeline (Research + Primary + 3 cross-checks) | $0.47 | - | $0.47 |
| 3 specialist panel additions (Reasoning-Verification, Execution-Feasibility, Market-Reality) | $0.21 | - | $0.68 |
| Market simulation engine | $0.06 | - | $0.74 |
| Vertical red-team pass | $0.05 | - | $0.79 |
| Corpus write, embedding, similarity search | $0.02 | (dual-score schema, same query cost) | $0.81 |
| Prediction tracking cron (amortized) | $0.01 | - | $0.82 |
| Infrastructure (sanitization gate, PDF automation, storage) | $0.04 | (extended gate coverage, same infra cost) | $0.86 |
| v3 fully-loaded cost/review | $0.86 | ||
| Dual scoring + per-dimension explanation generation | - | +$0.04 | $0.90 |
| Fix-It action plan generation (Pro+ only) | - | +$0.05 | $0.95 |
| Second-Opinion Audit Agent (Enterprise only, amortized across all reviews) | - | +$0.03 | $0.98 |
| v3.6 fully-loaded cost/review | $0.98 |
Chat-to-Refine and URL-to-Review add ~$0.02-0.03/session, treated as funnel/acquisition cost (not review COGS) - consistent with the free-tier loss-leader model. Free-tier full reviews remain ~$0.07/review; the lightweight Free-tier fix-it summary adds under $0.01.
| Tier | Price | Included Reviews/mo | Blended Rev/Review | Cost/Review | Gross Margin |
|---|---|---|---|---|---|
| Free | $0/mo | 1 | - (loss leader) | ~$0.07 | n/a |
| Pro | $79/mo | 20 | $3.95 | $0.98 | ~75% |
| Enterprise | $499/mo | 100 | $4.99 | ~$1.20 (incl. corpus/API/audit-agent infra) | ~76% |
| White-Label | $1,999+/mo | fair-use unlimited | volume-dependent | ~$1.20-1.35 | ~72-79% |
| Component | Model calls added per review | Notes |
|---|---|---|
| Chat-to-Refine | N (per chat turn, pre-review, not per review) | billed as acquisition cost, not COGS |
| URL Extractor | 0 model calls (extraction is deterministic scraping) + sanitization pass | negligible marginal cost |
| Dual Scorer | 0 (tagging happens inside existing Phase 2 call) | zero marginal model cost |
| Explainer | 0 (explanation text generated inline in Phase 2 call, longer output token count) | cost shows up as slightly higher Phase 2 token spend, captured in the $0.04 line |
| Fix-It Generator | 1 additional call, post-pipeline | $0.05/review, Pro+ only |
| Audit Agent | 1 additional call, Phase 6 | $0.03/review amortized, full cost when isolated to Enterprise-only billing is higher per Enterprise review |
| Rules Engine | 0 additional model calls (injected into existing Phase 2 prompt + deterministic pre-check) | cost is engineering/validation overhead, not inference |
| White-Label Engine | 0 model calls | pure infra/config cost |
| Community Layer | 0 model calls (human-generated) | moderation queue has infra cost, not inference cost |
Mapped to the v3.6 5-phase roadmap from the business proposal.
idea_score/proposal_score columns built in from day one (no later migration).org_id backfilled as NULL on existing rows.chat-api, chat-ws, chat_sessions/chat_messages) free-tier-wide.url-extractor worker, POST /review/url) alongside file upload.dimensions.score_type, Phase 5 aggregation logic).dimension_explanations, Phase 2 mandatory-field enforcement).remediation_plans, fixit-generator worker).audit_findings, Phase 6 worker, not yet billing-gating).rule_templates, org_rules, pre-check + Phase 2 injection logic).org_configs, domain-verifier, portfolio console) - capped at 3 pilot accounts (proposal Condition 4) until segmentation/confidentiality independently verified.peer_reviews, moderation-queue) - not enabled by default for any tier (proposal Condition 7).| Component | Failure Mode | Detection | Behavior / Recovery |
|---|---|---|---|
| Chat-to-Refine | Coach model produces full-paragraph ghostwritten content instead of coaching questions | Output length/completeness heuristic flags session for spot-check; periodic manual transcript audit (quarterly minimum) | Session flagged, not blocked in real time (would break UX); pattern-level drift triggers prompt-tuning pass, not per-session blocking |
| Chat-to-Refine | chat-ws socket drops mid-conversation |
Client-side reconnect with session resume via session_id |
REST polling fallback (POST /chat/{id}/message) available if WebSocket infra is degraded |
| URL Extractor | Target page is JS-rendered SPA Crawl4AI can't fully execute | Extraction returns near-empty content below a minimum-length threshold | 422 with reason code extraction_insufficient; user prompted to upload file instead |
| URL Extractor | Extraction exceeds timeout (45s hard cap) | Worker timeout | 422 extraction_timeout; no partial pipeline entry, no quota consumed |
| URL Extractor | Extracted content contains injection payload ("ignore previous instructions...") | Sanitization/injection-detection pre-pass on extracted text | Content flagged and either stripped of the payload or the review is rejected with 422 content_flagged, never silently passed through |
| Dual Scorer | Judges disagree on score_type tagging for the same dimension |
Audit Agent's tagging-consistency check (Phase 6) | Logged as tagging_inconsistency finding; Phase 5 aggregation falls back to the majority tag among judges, review still completes |
| Explainer | Primary Reviewer omits explanation.text or cited_gaps for a dimension |
Schema validation on Phase 2 output (fail-closed) | Phase 2 call is retried once with an explicit "missing required field" repair prompt; second failure escalates to human review queue rather than shipping an incomplete verdict |
| Explainer | Explanation text leaks internal architecture/vendor detail | Sanitization gate scan | Verdict withheld from egress until re-generated and re-scanned; review shows status=held_for_review to the user, not a silent partial response |
| Fix-It Generator | Remediation plan generation call fails or times out | Worker-level retry (3x exponential backoff) | On exhaustion, verdict ships without remediation; a background job retries generation and notifies the user when the plan is ready (does not block verdict delivery) |
| Fix-It Generator | Action items reveal too much evaluation logic (scoring-weight leakage) | Manual spot-check + pattern detection on re-review score deltas (superficial compliance vs. genuine improvement) | Prompt-level fix: regenerate remediation guidance templates to describe gaps, not weights; feed signal into reviewer-accuracy scoring per the proposal's mitigation |
| Audit Agent | Shadow-mode audit never files a finding across many reviews | Reviewer-accuracy-style tracking on the audit agent itself (judges.is_audit_agent=true) |
Treated as a signal the audit agent may be too conservative or echoing the primary panel; triggers prompt/config review before going live for billing |
| Audit Agent | Audit call fails/times out | Worker retry, then graceful degradation | Verdict ships without Phase 6 findings (audit.findings=[], status still done, not blocked) - audit is additive, never a hard gate on verdict delivery |
| Rules Engine | Org submits malformed rule JSON | JSON Schema validation against rule_templates.schema_json at save time |
422 with specific schema violation path; rule never persisted as is_active |
| Rules Engine | Hard-block rule incorrectly blocks a valid review (false positive) | Org admin dashboard shows block reason + rule name | Org admin can deactivate/edit the specific rule (org_rules.is_active=false) without needing engineering support |
| Rules Engine | Rule text attempts prompt injection via crafted field values | Constrained AST parsing + fixed template rendering (Section 7.3) | Injection payload is inert - it's substituted into a whitelisted template slot, never concatenated as raw instruction text |
| White-Label | Custom domain CNAME/TXT never verifies | domain-verifier polling times out after 48h |
Org notified with specific DNS record diagnostics; domain stays unverified, org continues on default subdomain until resolved |
| White-Label | Cross-tenant data leak (application bug omits org_id filter) |
RLS policy at database layer blocks the query regardless of application bug | Query returns empty/denied rather than cross-tenant data; incident logged via audit_log, alerts on-call |
| White-Label | Branding CSS injection via brand_colors/logo_url fields |
Input validation (hex color regex, URL allowlist for logo hosting) | Malformed values rejected at org-api before persisting to org_configs |
| Community Layer | Roast/Boost comment contains harassment/abuse | User-driven abuse reporting (peer_review/{id}/report) β moderation queue |
Comment flagged (moderation_status='pending_review'), hidden from submitter view pending moderator action; repeated flags trigger user-level cooldown |
| Community Layer | Rate limit bypass attempt (rapid submissions) | Redis sliding-window counter + Postgres durable count, double-enforced | Requests past the 3/day cap return 429; discrepancy between Redis and Postgres counts triggers a reconciliation job, not silent over-limit acceptance |
| Community Layer | Feature accidentally enabled by default for a tier | Feature flag (roast_boost_enabled) is off by default at the flag-service level, not the application code level |
Flag flip is a config change, auditable and instantly reversible without a deploy |
| Sanitization Gate (all surfaces) | Gate itself fails/errors on a scan | Fail-closed design: no scan result = no egress | Content held, status=held_for_review; alerts on-call immediately since this blocks all outbound traffic for that review, by design |
| Cross-cutting | Any Phase 1-6 model call fails | Existing v3 degraded-mode fallback | Pipeline runs with fewer judges rather than failing outright (v3 policy, unchanged in v3.6); dual-score/explanation/audit outputs degrade gracefully to "insufficient panel data" rather than fabricating scores |
| Architecture Section | Proposal Reference |
|---|---|
| Β§5.1 Chat-to-Refine | Proposal Β§05, Feature 1 |
| Β§5.2 URL Extractor | Proposal Β§05, Feature 2 |
| Β§5.3 Dual Scorer | Proposal Β§06, Feature 3 |
| Β§5.4 Explainer | Proposal Β§06, Feature 4 |
| Β§5.5 Fix-It Generator | Proposal Β§07, Feature 5 |
| Β§5.6 Audit Agent | Proposal Β§08, Feature 6 |
| Β§5.7 Rules Engine | Proposal Β§08, Feature 7 |
| Β§5.8 White-Label Engine | Proposal Β§09, Feature 8 |
| Β§5.9 Community Layer | Proposal Β§09, Feature 9 |
| Β§9 Implementation Sequence | Proposal Β§12, Phases 0-4 |
| Β§8 Cost Model | Proposal Β§14, Financial Model |
| Β§7.2 Sanitization Gate | Proposal Β§16, Condition 1 |
| Β§7.4 Multi-Tenant Isolation | Proposal Β§16, Condition 4 |
End of document.