v3 gave the pipeline a memory. v3.5 gives the user a relationship with it -- a coach before submission, a second score that separates the idea from the pitch, a plain-language explanation of every score, a fix-it plan after, and a way for real people and real partners to plug into the loop. Nine new capabilities, sourced from a structured competitive landscape scan, merged into the existing v3 roadmap rather than bolted on top of it.
VerdictTank v3 solved memory: the pipeline now remembers every review it has ever run and checks whether its own predictions came true. That work stands. v3.5 does not replace it -- it wraps it in a product journey that starts before a user submits anything and continues after the verdict lands.
The trigger for v3.5 is a structured scan of the competitive landscape (Competitive Landscape Research Batch 001), covering adjacent categories: pre-submission coaching tools, URL-based intake products, community-scored idea platforms, audit-loop review patterns, no-code enterprise rules engines, and monitoring-plus-remediation SaaS. Seven concrete features and two cross-cutting strategic themes emerged as directly applicable to VerdictTank without diluting its core identity as a brutally honest, multi-judge critique engine.
Two themes recur across nearly every competitor studied and deserve top billing:
Theme 3, Transparency Sells: the products that explain why a score is what it is retain users at a materially higher rate than products that hand down a number and walk away.
Theme 5, The Fix-It Layer: the products that turn a one-shot assessment into a repeatable loop (assess, then help fix, then re-assess) convert a single transaction into a subscription relationship.
v3.5 integrates nine new capabilities directly into the v3 structure, spanning three product moments: before the review (Pre-Review Coach), during the review (Transparency Layer, Trust & Enterprise Controls), and after the review (Fix-It Layer, Distribution Expansion). None of these replace a v3 enhancement -- all nine sit alongside the existing corpus, prediction tracking, vertical templates, and API work, and several of them depend directly on v3 infrastructure already scoped (the sanitization gate now has more surfaces to cover; the corpus now stores two scores per review instead of one).
7 features from direct competitive research findings, 2 from cross-cutting strategic themes. All merged into existing sections, not appended.
The Fix-It Layer is the single highest-leverage change: reviewing stays paid, fixing is included, and users come back to re-score after acting on it.
Coach before, dual transparent score during, remediation plan and community signal after. VerdictTank stops being a single checkpoint and becomes a loop.
Two layers of baseline matter for this proposal: the live v2 platform, and the v3 upgrade already scoped and conditionally approved on top of it.
Four cross-vendor reviewers today, ~$0.47/review fully loaded, 3-9 minute turnaround, tiered $19-299/mo pricing. No memory of past reviews, no vertical specialization, no API, no public-facing output, and sanitization/PDF generation both handled manually.
v3 (Section 3) is the conditionally-approved fix for the memory and infrastructure gaps: a review corpus, prediction-vs-outcome tracking, reviewer accuracy scoring, vertical templates, adversarial red-teaming, market simulation, shareable reports, an API, and the sanitization/PDF automation everything else depends on.
What v2 and v3 both still lack, and what surfaced repeatedly in the Batch 001 competitive scan: nothing helps the user before they submit, nothing explains a score in plain language, nothing tells the user what to do about a bad score, and nothing lets outside partners (consultants, accelerators) run the product as their own without being folded into the same generic White-Label tier as everyone else. That is the gap v3.5 closes.
Everything approved in v3 remains in scope for v3.5. This section is a compressed recap so the rest of this document reads as an integration, not a rewrite. Full detail lives in the v3 proposal.
Every review result, searchable: scores, flaws, conditions, predictions.
T+90/180/365 cron checks whether flagged risks actually materialized.
Judges earn a track record; future weighting follows the data.
Auto-classified by vertical, reviewed against industry benchmarks.
Industry-specific attack vectors: regulatory, compliance, churn dynamics.
12-month simulated trajectory replaces the static financial table.
Sanitized, shareable score cards and summaries, opt-in.
Full pipeline as an endpoint. Upload a document, get a structured review.
Every new proposal benchmarked against the corpus, by vertical.
Pre-deploy gate strips model/architecture details before anything goes public.
Scripted end-to-end report generation, required for API and Enterprise volume.
Nine new capabilities, sourced from Competitive Landscape Research Batch 001: seven direct feature findings (labeled by priority as scanned) and two cross-cutting strategic themes. Each is placed into the product moment it belongs to, not treated as a standalone bucket.
| # | Capability | Source Pattern | Priority | Product Moment | Timeline |
|---|---|---|---|---|---|
| 1 | Chat-to-Refine | Pre-submission AI coaching pattern | P1 | Before review | 1-2 wks |
| 2 | URL-to-Review | URL-based content intake pattern | P1 | Before review | 1-2 wks |
| 3 | Dual Scoring (Idea + Proposal) | Split-signal scoring pattern | P1 | During review | 2-3 wks |
| 4 | "Explain the Low Score" transparency | Strategic theme: Transparency Sells | P1 | During review | 2 wks |
| 5 | "Here's How to Fix It" remediation | Strategic theme: The Fix-It Layer | P2 | After review | 4-5 wks |
| 6 | Second-Opinion Audit Agent | Audit-loop review pattern | P2 | During review | 6-8 wks |
| 7 | Configurable Review Rules Engine | No-code enterprise rules pattern | P2 | Enterprise controls | 3-4 wks |
| 8 | White-Label for Consultants (dedicated track) | Branded-service reseller pattern | P2 | Distribution | 2-3 wks |
| 9 | Roast/Boost Community Peer Review | Community scoring pattern | P3 | After review | 3-4 wks |
Read together: items 1-2 lower the barrier to a good submission, items 3-4 make the resulting score legible and trustworthy, item 5 makes the score actionable and brings users back, items 6-7 harden the product for enterprises with real compliance and quality-control needs, and items 8-9 expand who VerdictTank reaches without diluting the core judgment engine.
Two new capabilities sit between the user and the pipeline, lowering the barrier to a submission worth reviewing at all.
Before submission, an AI chat helps the user structure their argument, surface gaps, and tighten clarity. Not a reviewer -- a coach. It never scores; it only helps the user get their best draft in front of the pipeline. Timeline: 1-2 weeks.
Paste a URL instead of uploading a file. The system extracts the content, structures it into the standard intake format, and runs it through the same pipeline. Removes friction for pitch decks published as web pages, Notion docs, or landing pages. Timeline: 1-2 weeks.
The chat-to-refine pattern (seen in pre-submission coaching tools in the competitive scan) works because it is decoupled from judgment. If the coach and the judge were the same conversation, users would learn to argue with the score instead of improving the proposal. Keeping the coach free-standing and pre-submission preserves the brutal-honesty brand of the actual review while still lowering the failure rate of first-time submissions.
Chat-to-Refine is offered as a free, unlimited-use funnel step available to every tier including Free -- it costs a fraction of a full review and its entire purpose is to produce more, and better, paid reviews downstream. URL-to-Review is available at every paid tier and functions as an alternate intake path alongside file upload, not a separate product.
Three capabilities make the verdict itself more legible: two new scores instead of one, and a plain-language explanation behind every dimension score.
Every review now produces two numbers, not one: an Idea Score (0-100, is the underlying concept sound) and a Proposal Score (0-100, how well was it argued). A brilliant idea poorly pitched and a weak idea beautifully argued now produce visibly different signals instead of collapsing into one composite number.
Every per-dimension score ships with a specific, evidence-backed explanation. Not "Market Analysis: 4/10" -- "Market Analysis: 4/10 -- no TAM calculation, no competitor pricing data, assumes zero competition." Direct application of the transparency pattern seen across the strongest-retaining products in the scan.
This is the feature most likely to be described by users, unprompted, as "the thing nobody else does." Single-score critique tools conflate concept quality with execution quality, which means a founder with a great idea and a mediocre deck gets the same verdict as a founder with a mediocre idea and a great deck. VerdictTank v3.5 tells them apart. It also changes what the corpus can say: "your Idea Score is in the 80th percentile but your Proposal Score is in the 30th" is a sharper, more actionable comparison than a single blended percentile.
The existing 10-dimension rubric splits cleanly: dimensions like market size, competitive moat, and unit economics feed the Idea Score; dimensions like clarity of argument, evidence quality, and internal consistency feed the Proposal Score. The corpus database (v3, Section 3) stores both scores per review from day one rather than retrofitting later -- this is the one place where a v3.5 feature changes a v3 schema before it ships. The per-dimension explanation text (item 4) is generated as a required field on every dimension score, not an optional add-on, and is what the sanitization gate (v3 Infrastructure Fix #10) must additionally scan before anything reaches a public report or API response, since explanation text is exactly the kind of freeform output most likely to leak internal architecture detail if left unchecked.
After the verdict, the system generates a structured action plan tied to the low-scoring dimensions: "To improve your Market Analysis score: (1) calculate TAM using this formula, (2) add a competitor pricing table (template provided), (3) cite 3 industry reports." The submitter can then revise and re-submit.
Every AI critique product on the market, VerdictTank v2 and v3 included, is structurally a one-shot transaction: pay, get scored, done. The Fix-It Layer converts that into a loop. The competitive pattern here (seen in monitoring-plus-remediation SaaS in the scan) inverts the usual model where monitoring/detection is free and fixing is the paid upsell -- VerdictTank flips it deliberately: reviewing is paid, fixing is included. That is the differentiator, and it is the reason a Pro or Enterprise subscriber comes back inside the same billing month instead of churning after one review.
Combined with dual scoring (Section 6) and re-review, this also produces the cleanest possible proof point for reviewer accuracy: a user who follows the fix-it plan and re-submits generates a natural before/after score delta that the corpus can track and eventually surface as a product stat ("proposals that follow the Fix-It plan improve their Proposal Score by a median of X points on re-review").
Available on Pro and above. Re-review after applying a Fix-It plan is billed as a normal review against the plan's monthly allotment (or overage rate) -- it is not a free re-run, since the AI cost of a full pipeline pass is the same either way. What is "included" is the remediation plan itself, not unlimited re-scoring. Free tier gets a lightweight, single-paragraph fix-it summary as a taste of the feature; the full structured, dimension-by-dimension action plan with templates is a paid-tier feature.
Two capabilities target Enterprise and White-Label customers who need the panel's judgment checked and its criteria tailored to their own standards.
After the primary multi-judge review completes, a separate audit agent reviews the review itself -- checking for blind spots, groupthink, or dimensions the panel underweighted. Applies an audit-loop pattern seen in adjacent review-automation tooling directly to proposal review. Timeline: 6-8 weeks.
Enterprise and White-Label customers define custom review rules through a no-code builder: industry-specific compliance checks, internal evaluation criteria, brand voice guidelines. The panel's standard rubric still runs; custom rules layer on top rather than replacing it. Timeline: 3-4 weeks.
Both features exist to answer the same institutional-buyer objection: "how do I trust an AI panel's judgment enough to put my firm's name behind it, or bend it to our own criteria?" The audit agent answers the trust question with a second independent check baked into the pipeline. The rules engine answers the customization question without opening the door to arbitrary prompt injection into the panel -- custom rules are additive checks evaluated alongside the fixed rubric, never a replacement for it, which keeps VerdictTank's brutal-honesty core intact even when White-Label and Enterprise customers bring their own criteria.
The audit agent's own findings are subject to the same reviewer-accuracy tracking (v3, Tier 1) as the primary panel -- an audit agent that never finds anything, or that only echoes the primary panel, is itself a signal the accuracy scoring system should catch.
The v3 White-Label tier (Section 11) already covers accelerators, VCs, and consultancies at the pricing level. v3.5 makes it a dedicated product track rather than a checkbox on the Enterprise tier: custom domains, custom logos, branded email templates, and a consultant-facing admin console for managing client portfolios. Consulting firms and accelerators can now offer VerdictTank as their own branded review service end to end. Timeline: 2-3 weeks.
An optional, togglable community layer: peers can leave "roast" (critical) or "boost" (supportive) signals on a proposal. This is a supplementary human signal, not a replacement for the AI panel, and is off by default -- the submitter opts in per-proposal. Timeline: 3-4 weeks.
Item 8 does not create a new pricing tier; it defines what "White-Label" already means in v3 pricing more precisely, with concrete deliverables (custom domain, custom branding, portfolio console) instead of a general promise. Item 8 is the reason the White-Label tier's feature list is expanded in Section 11 without changing its price.
Item 9 is intentionally P3 and opt-in. Community layers carry real risk (brigading, low-signal noise, moderation overhead) that does not belong anywhere near the core judgment product by default. It ships last, behind a feature flag, and is evaluated after one full quarter of opt-in data before any decision is made about turning it on by default for any tier.
The v3 panel expansion from 4 to 7 reviewers (Reasoning-Verification, Execution-Feasibility, Market-Reality judges) is unchanged and carried forward in full -- see the v3 proposal for the complete rationale and the deferred-candidate table. v3.5 adds one new role to the panel's output contract rather than its roster:
No new judge is added for dual scoring (Section 6). Instead, each existing judge's dimension scores are tagged at generation time as contributing to the Idea Score or the Proposal Score, and the Second-Opinion Audit Agent (Section 8) additionally checks that this tagging is applied consistently across judges before a verdict is finalized. This keeps the panel at 7 reviewers plus the audit agent, rather than growing it further, consistent with the v3 condition that panel size is a tunable parameter, not a one-way ratchet.
Per house sanitization policy, exact vendor and model identities remain internal and are stripped from all public-facing documentation and API responses by the automated sanitization gate, exactly as under v3.
v3.5 keeps the v3 tier structure and price points. Four features change what a tier includes; none change the sticker price. The current verdicttank.com beta pricing ($29/mo "Inner Circle") is treated as a temporary bridge tier that sunsets at v3.5 General Availability.
Pay-per-review outside a subscription: $19.99, full pipeline, one-time. Annual billing on Pro and Enterprise: 20% discount. All tiers retain the degraded-mode fallback from v2/v3 -- if a judge is unavailable, the pipeline runs with fewer judges rather than failing.
The v3 roadmap's four phases are re-sequenced into five phases to absorb the nine v3.5 items at the point in the pipeline where their dependencies are actually satisfied. Phase 0 now covers sanitization broadly enough for both v3 and v3.5 public surfaces, since patching it twice would be wasted work.
| Phase | Timeline | Scope | Exit Criteria |
|---|---|---|---|
| Phase 0 Foundation |
Weeks 1-3 | Automated sanitization gate live and blocking deploys, scoped to cover free-text explanation fields (v3.5 item 4) in addition to standard report/API surfaces. Automated PDF generation scripted end-to-end. Corpus database schema designed and backfilled with every historical review, with dual-score columns (Idea Score / Proposal Score) built into the schema from the start rather than retrofitted. | Zero manual sanitization steps remain. Every past review is queryable in the corpus. Schema supports dual scoring without migration. |
| Phase 1 Coach & Intake |
Weeks 4-6 | New: Chat-to-Refine ships free-tier-wide. URL-to-Review intake ships alongside file upload. Corpus search live internally. Prediction-vs-outcome cron scheduled at T+90/180/365. Reviewer accuracy scoring begins tracking (read-only). | Chat-to-Refine and URL-to-Review both in production with zero sanitization gate failures. First cron cycle completes on a batch of 90-day-old reviews with real outcome data attached. |
| Phase 2 Transparency & Depth |
Weeks 7-12 | Dual Scoring (Idea Score + Proposal Score) ships as the panel's default output contract. "Explain the Low Score" per-dimension transparency ships alongside it. Vertical auto-classification and templates for the top 4 verticals by volume. Adversarial red-team per vertical. Market simulation engine replaces the static financial table. | Every review issued after Phase 2 exit carries two scores and a per-dimension explanation. 3 of 4 vertical templates validated against known historical outcomes. |
| Phase 3 Fix-It & Trust |
Weeks 13-19 | "Here's How to Fix It" structured remediation plan ships on Pro and above. New reviewer panel additions (Section 10) go live. Second-Opinion Audit Agent begins auditing Enterprise-tier reviews. Configurable Review Rules Engine (no-code) opens to Enterprise/White-Label in closed beta. | Fix-It plans generated on 100% of Pro+ reviews. Audit Agent completes 90 days of shadow-mode review with a documented blind-spot catch rate before going live for billing purposes. |
| Phase 4 Distribution & GA |
Weeks 20-25 | Public-facing shareable review reports ship. Review-as-a-Service API opens in closed beta. Competitive corpus comparisons appear on every review. White-Label for Consultants dedicated track (custom domain, logo, portfolio console) launches. Roast/Boost community layer ships behind an opt-in flag. Reviewer accuracy scoring begins weighting live verdicts. API opens to Enterprise tier generally. Beta pricing sunsets; existing beta accounts enter the grandfathered migration window. | API beta completes 100 reviews with zero sanitization gate failures. First White-Label-for-Consultants pilot renews past month one. Beta migration communicated to 100% of existing beta accounts with a 6-month grandfather window confirmed. |
| Category | Players | What They Do | VerdictTank v3.5 Difference |
|---|---|---|---|
| AI proposal writers | Template-based generation tools | Generate proposals from prompts | We critique them; we still don't write them -- but Chat-to-Refine now helps a user structure their own argument before submission, closing the gap without becoming a generation tool ourselves. |
| Pre-submission coaching tools | AI chat-based drafting assistants | Help structure an argument before submission, no independent judgment afterward | Chat-to-Refine matches this pattern for intake, then hands off to an independent, brutally honest multi-judge panel these tools don't have. |
| Single-score idea-validation platforms | Community or single-model scoring tools | One blended score, idea quality and pitch quality conflated | Dual Scoring (Idea Score + Proposal Score) separates the two signals -- the single differentiator most competitors in this category lack entirely. |
| Human consultants | Independent reviewers, boutique firms | Manual review, $300-800/hr, days of turnaround | Under $1 and minutes, now with a Fix-It plan included -- consultants charge extra for that follow-up work; we do not. |
| Monitoring/audit SaaS with paid remediation | Detect-then-upsell-the-fix platforms | Monitoring or detection free, remediation is the paid tier | VerdictTank inverts this: reviewing is paid, fixing is included. Removes the friction point that causes users to detect a problem and never pay to fix it. |
| Accelerator/VC diligence tools | Internal scoring rubrics, deal-flow CRMs | Human-scored, not benchmarked against a broader corpus | White-Label for Consultants gives accelerators a benchmarked, auditable, corpus-backed second opinion with their own branding, domain, and portfolio console -- not just a shared login. |
| Cumulative-memory, full-journey proposal review | No direct competitor exists as of August 2026. A platform that remembers its own track record, tells two scores instead of one, explains every dimension in plain language, and hands the user a fix-it plan is a genuinely new category composition, not an incremental feature. | We are still the category, now with the full loop: coach, judge, explain, fix, re-score. |
| Cost Component | Per Review |
|---|---|
| Base pipeline (research + primary + 3 cross-checks) | $0.47 |
| 3 new specialist reviewer additions (v3, Section 10) | $0.21 |
| Market simulation engine | $0.06 |
| Vertical red-team pass | $0.05 |
| Corpus write, embedding, and similarity search (dual-score schema) | $0.02 |
| Prediction tracking cron (amortized) | $0.01 |
| Infrastructure (sanitization gate, PDF automation, storage) | $0.04 |
| New: Dual scoring + per-dimension explanation generation | $0.04 |
| New: Fix-It action plan generation (Pro+ only) | $0.05 |
| New: Second-Opinion Audit Agent (Enterprise only, amortized across all reviews) | $0.03 |
| Fully loaded cost per full v3.5 review | $0.98 |
Chat-to-Refine and URL-to-Review add negligible marginal cost (~$0.02-0.03 per session) and are treated as funnel/acquisition cost, not review COGS, consistent with the free-tier loss-leader model already established in v3. Free-tier full reviews stay near $0.07/review; the lightweight fix-it summary on Free adds under a cent.
| Tier | Included Reviews/mo | Blended Revenue/Review | Cost/Review | Gross Margin |
|---|---|---|---|---|
| Pro ($79/mo) | 20 | $3.95 | $0.98 | 75% |
| Enterprise ($499/mo) | 100 | $4.99 | ~$1.20 (incl. corpus/API/audit-agent infra) | 76% |
| White-Label ($1,999+/mo) | Fair-use unlimited | Volume-dependent | ~$1.20-1.35 | 72-79% |
| Year 1 | Year 2 | Year 3 | |
|---|---|---|---|
| Paying users (Free excluded) | 210 | 780 | 2,450 |
| White-label / consultant-track accounts | 5 | 18 | 48 |
| MRR (year-end) | $13,100 | $54,000 | $176,000 |
| Annual revenue | $88,000 | $412,000 | $1,470,000 |
| AI + infra costs | $15,200 | $59,000 | $188,000 |
| Net | $72,800 | $353,000 | $1,282,000 |
Uplift over v3's original forecast ($52.2K / $277K / $1.098M net) comes primarily from three effects: Chat-to-Refine and URL-to-Review lowering intake friction and increasing free-to-paid conversion; the Fix-It Layer converting one-shot reviews into repeat-use within the same billing month, raising blended reviews-per-paying-user; and the dedicated White-Label-for-Consultants track making that tier easier to sell against accelerator-specific budget lines than a generic enterprise pitch. Assumptions carried forward from v3: AI cost deflation of 15-20% annually, 5-7% monthly Pro churn, materially lower Enterprise/White-Label churn given contractual terms. The beta-to-Pro migration (Section 11) is modeled as net-neutral to Year 1 revenue -- beta accounts are assumed to already be counted in the Year 1 paying-user base at their eventual Pro-equivalent value once the grandfather window closes.
The eight risks identified in v3 remain in force unchanged and are carried forward in full (corpus confidentiality, accuracy-scoring convergence, prediction attribution difficulty, sanitization gate failure, API abuse, panel cost creep, competitive-comparison leakage, white-label brand risk, training-data recursion). v3.5 adds five new risks introduced by the nine new capabilities.
| Risk | Likelihood | Impact | Mitigation |
|---|---|---|---|
| Chat-to-Refine coaching drifts into ghostwriting the proposal for the user, undermining the "we critique, we don't write" positioning | Medium | High | Coach is scoped to structuring questions and gap-identification prompts only, never full-paragraph generation. Coach transcripts are logged and periodically spot-checked for drift toward ghostwriting behavior. |
| Dual scoring is perceived as a gimmick or confuses users if the two scores are not clearly differentiated in the UI/report | Medium | Medium | Every report leads with a two-line plain-language framing ("your idea is strong, your pitch needs work" style) before showing the numeric breakdown. User-test the framing before Phase 2 exit. |
| Per-dimension "Explain the Low Score" freeform text becomes the highest-risk sanitization surface -- more opportunity for an internal detail to leak than a fixed-template report ever had | Medium | High | Sanitization gate (v3 Infrastructure Fix #10) explicitly extended in Phase 0 to scan all freeform explanation and remediation text, not just fixed report sections. Manual spot check on the first 200 explanation outputs before Phase 2 GA, not just the first 50 API responses as in v3. |
| Fix-It remediation plans give away enough of the pipeline's evaluation logic that a sophisticated user reverse-engineers what the panel is checking for, gaming future submissions rather than genuinely improving them | Medium | Medium | Remediation plans stay tied to what is missing (data, evidence, structure), not to how the panel weighs it. Track re-review score deltas for a pattern of superficial compliance (template-filling without substance) versus genuine improvement, and feed that signal back into reviewer accuracy scoring. |
| Roast/Boost community layer attracts low-signal noise, brigading, or becomes a vector for harassment tied to a submitter's real proposal | Medium | Medium | Opt-in only, off by default, no default visibility to anyone outside the submitter's own dashboard unless the submitter explicitly shares it. Ships behind a feature flag in Phase 4 and is evaluated after one full quarter of opt-in data before any default-on consideration. |
v3.5 is an integration, not a rewrite -- it takes the v3 memory foundation as given and wraps it in a full product journey (coach, dual score, transparency, fix-it, re-score) that the competitive scan shows nobody else in this category has assembled end to end. The two strategic themes it centers, transparency and the fix-it loop, are the two changes most likely to move retention and repeat-use, which is where v3 alone was weakest. All four v3 conditions carry forward unchanged; three new conditions are added specific to the new surfaces.
This proposal will be submitted to the VerdictTank pipeline for its own review before implementation begins, consistent with house policy that every IT Pro Partner product proposal is reviewed by the tool it is proposing to improve.