v5.0 · Pre-Revenue · Validation-Tested

14 blind errors your solo model missed

A solo frontier model gives you a smooth, confident score. An 11-judge panel gives you the 14 things it was wrong about. VerdictTank does not sell you a higher number. It sells you the errors that number was hiding.

Category: AI Proposal Review & Error Detection Panel: 11 seats · 9 vendors Architecture: Full technical document Judge pool: Spec v2.3

01The Thesis Changed, Because the Data Said So

VerdictTank v4.0 was sold on a claim we could not defend: that a multimodel panel produces a better score than a single strong model. On 2026-08-12 we ran that claim against real proposals and it failed. What we found instead is a stronger product.

14
Material errors caught
-0.40
Aggregate score delta
3
Real proposals scored
11
Judge seats, 9 vendors

What failed

The score-elevation thesis. Across three real proposals the panel mean came in 0.40 points below the solo baseline. The panel did not lift scores. On two of three proposals it pushed them down. We are publishing that result rather than burying it, because the reason it happened is the product.

What worked

Error detection density. The same panel run surfaced 14 material errors that the solo baseline missed or underweighted: revenue arithmetic that was wrong by a factor of seven, a funded direct competitor the solo pass never named, a launch-blocking compliance cost larger than projected first-year revenue. None of those show up as a score. All of them decide whether the proposal wins.

You do not buy VerdictTank to get a higher score.
You buy it to find the $60K compliance hole and the broken contractor budget before you ship.

Why the spread is the signal

A single model scoring alone produces low variance. It reads the document once, forms one coherent opinion, and every dimension it emits is downstream of that opinion. The result feels authoritative precisely because nothing inside it disagrees.

An 11-seat panel of nine different vendors cannot produce that coherence, and the incoherence is diagnostic. When a Financial Integrity judge scores a proposal 8.2 while an Execution Feasibility judge scores the same document 2.8, that 5.4-point spread is not noise. It is a precise statement: the money works, the delivery plan does not. A solo model averages that tension away into a single confident 6.1 and tells you nothing actionable.

Measured, not asserted: on every one of the three validated proposals, panel spread exceeded solo spread. RFP Tank 3.7 vs 1.4. VentureBuilt 5.4 vs 2.4. CartMySupply 2.6 vs 1.8. Widening variance is the intended behavior, not a defect to tune out.

02How We Found This: Four Generations and a Self-Review

VerdictTank is a proposal review engine, not a proposal writer. It ingests a finished document and returns scored dimensions plus a ranked list of concrete Fix-It items. The pipeline was not designed in the abstract. It was hardened across four architectural generations, and then it was pointed at its own proposals.

v2

Single-Model Scorer

Breakthrough: the dual-axis rubric

v2 established the core insight: a proposal has two independent quality axes. Narrative quality (clarity, structure, persuasion) and compliance quality (does it actually answer the scored requirements). One model scored both from one prompt. It proved the concept and exposed the flaw: the axes bled together. A beautifully written section that missed a mandatory requirement scored too high, because the same reasoning pass that admired the prose also graded the compliance.

v3

Separated Scoring Passes

Breakthrough: axis isolation

v3 split scoring into two independent passes with two purpose-built prompts. The narrative pass never sees the compliance rubric. The compliance pass never rewards eloquence. This is the decision that makes the dual score trustworthy: the two numbers can now disagree, and their disagreement carries information. A 9/10 narrative next to a 4/10 compliance is a proposal about to lose.

v4

Multimodel Adversarial Review

Breakthrough: cross-model verification

A single model scoring in isolation is confidently wrong at a predictable rate. v4 introduced a multimodel pipeline: a fast model produces first-pass scores and Fix-It candidates, then a stronger model reviews that output adversarially, challenging every deduction and confirming each Fix-It maps to real proposal text. Scores stopped drifting between runs. This is the generation that made the output defensible.

v5 · Current

Specialist Panel and the Integrity Gate

Breakthrough: disagreement as output

v5 replaces the adversarial pair with an 11-seat specialist panel across nine vendors, and adds a synthesis seat whose only job is to compute panel statistics, flag scores more than 1.5 standard deviations from the mean, and reconcile the verdict against the evidence. The output is no longer a number. It is a number, a spread, an outlier list, and a ranked set of material errors with the judge that caught each one.

The self-review that broke the old thesis

Before selling a review engine we ran the engine on our own work. We assembled the panel and scored three real proposals, in full, with the same prompts and rubric a paying customer would get. One of the three was VerdictTank's own sibling product. The panel returned a NO GO on it.

That run cost roughly $150 in inference and returned 8 of 11 seats. Two seats were lost to a provider credit wall hit mid-run and one to a model family that could not be dispatched at all. The incomplete panel is why v2.3 of the judge pool spec now requires a pre-flight health gate and a pre-baked failover roster, covered in section 5. The results below are what those 8 seats produced, and we report them at 8 seats rather than extrapolating to 11.

The pipeline is its own reference implementation. The full architecture is documented in the companion technical architecture document to a standard where an engineer can implement it from the spec alone. VerdictTank is pre-revenue. We make zero claims about users, beta cohorts, or external validation. What we claim is narrower and verifiable: the architecture is built, the pipeline runs, it was executed against three real proposals on 2026-08-12, and it failed its own headline thesis in public.

03Validation Run: 3 Real Proposals, 8 Reporting Seats

Every figure in this section comes from the 2026-08-12 validation run. Nothing is modeled, projected, or illustrative. The solo baseline is Claude Opus 5 scoring the same documents against the same 10-dimension rubric.

Panel mean vs solo baseline

ProposalPanel meanSolo baselineDeltaPanel spreadSolo spreadVerdict
RFP Tank v1.04.404.93 -0.533.71.4 NO GO
VentureBuilt v26.146.10 +0.045.42.4 CONDITIONAL GO
CartMySupply4.295.00 -0.712.61.8 NO GO
Aggregate4.945.34 -0.40Panel spread exceeded solo spread on all 3 THESIS FAIL

Thesis under test: the panel must show a greater than 0.5 point advantage over the solo mean to justify premium pricing. Result: FAIL on all three proposals and FAIL in aggregate. Panel composition for this run was 8 reporting judges (4 Band A, 4 Band B) out of 11 specified seats, a 73% coverage rate.

Why the panel scored lower

The panel does not elevate scores. It sharpens error detection, and error detection on a flawed document moves the number down. All three proposals contained severe cross-cutting defects that additional specialist scrutiny exposed more precisely: fatal execution gaps, competitive mispositioning, and legal blockers. The solo baseline was directionally correct on all three. The panel added precision, not points.

That is the entire finding, and it inverts the sales pitch. If your proposal is sound, the panel will roughly agree with a good solo model and cost you more. If your proposal has a hole in it, the panel finds the hole and the solo model does not. You are not buying a score. You are buying the probability that a specific, expensive, named mistake gets caught before an evaluator or an investor finds it for you.

Specialist divergence, measured

ObservationEvidence from the runWhat it means
Generalist seats run optimistic Gemini Pro scored RFP Tank 6.7 as a Band A generalist and 3.0 as the Band B Market specialist. Same model, same document, 3.7 points apart. Band A generalist scoring without specialist cross-check is systematically over-optimistic. The role, not the model, drives the score.
Role divergence beats model divergence DeepSeek V4 Pro scored VentureBuilt 6.4 as Cross-Check C and 2.8 as Execution Feasibility. A 3.6 point split inside one vendor. Panel diversity is not primarily about buying different vendors. It is about buying different questions.
One seat can flip a verdict Remove the 2.8 Execution score from VentureBuilt and the panel averages 6.6, reading as a clean GO. With it, the verdict is CONDITIONAL GO with a named contractor-budget fix. The lowest score in the panel is frequently the only one doing work. Averaging is what a solo model already does.
Tight clustering is also a signal CartMySupply produced zero outliers beyond 1.5 sigma and the tightest spread of the three (sigma 0.89). Unanimity across nine vendors on a low score is a far stronger NO GO than one model's low score.

04The 14 Errors: Every One Named

This is the product. Fourteen material errors the 8-judge panel caught that the solo baseline missed or underweighted, grouped by failure class. Each is a real finding from the 2026-08-12 run against a real document.

3
Revenue arithmetic errors
Headline numbers that contradict the proposal's own inputs.
4
Competitive mispositionings
Named, funded, shipping incumbents the document treated as absent.
3
Execution infeasibilities
Build plans that cannot be delivered at the stated budget or timeline.
2
Legal compliance blockers
Registration and privacy obligations that gate launch entirely.
2
Team capacity impossibilities
Founder hour budgets that exceed the hours available.
14
Total, across 3 documents
Mean 4.7 material errors per proposal reviewed.

RFP Tank v1.0 · panel 4.40 vs solo 4.93

ClassError the panel caughtCaught by
RevenueThree mutually inconsistent Year-1 revenue figures inside one document: $1.2M, $1.361M, and $372K. Plus a 22% MRR ramp inconsistency the narrative never reconciles.Financial Integrity
CompetitiveCLEATUS is a real, funded competitor at $4M seed with public product-led pricing of $39 to $250/mo, occupying the identical quadrant. The proposal does not name it. The Band A generalist seat actually cited CLEATUS pricing as a positive signal.Market Reality
CompetitiveGovEagle pricing referenced at a 15x inconsistency against the proposal's own comparison table.Market Reality
ExecutionFive of seven features marked TO BUILD at HIGH effort. The real-time Compliance Copilot alone needs 2 to 3 developers for 8 to 12 weeks. The plan allocates 4 weeks, solo.Execution Feasibility
TeamA solo founder shipping a 7-feature AI SaaS in 10 weeks, with hiring contingent on revenue that requires the product to already exist. A closed loop with no entry point.Team / Founder
LegalNo privacy policy and no terms of service, against FAR and CUI exposure, with ITAR implications on German-hosted infrastructure.Legal / Regulatory

Panel verdict: NO GO. Estimated rework 40+ hours. Recommendation is to cut scope to two features, extend to 20 weeks, hire a second developer before month one, rebuild the financial model, and address CLEATUS directly.

VentureBuilt v2 · panel 6.14 vs solo 6.10

ClassError the panel caughtCaught by
ExecutionContractor budget broken by a factor of 4 to 7. The stated $1,500/mo implies $11 to $22 per hour against a market rate of $75 to $100. At real rates that budget buys 105 to 140 hours and leaves roughly 800 hours on the founder.Execution Feasibility
RevenueYear 2 stated on a run-rate basis rather than recognized revenue. Restated correctly, the healthy scenario loses roughly $11K to $18K.Financial Integrity
CompetitiveThe uniqueness claim is contradicted by shipping products. LivePlan Plan Review and IdeaProof already occupy the space.Market Reality
Team37 engagements plus 950 hours plus an MSP day job. The three commitments cannot coexist in one calendar.Team / Founder

Panel verdict: CONDITIONAL GO with six named conditions. Estimated rework 15 to 20 hours. This is the case that most clearly shows the value: the panel mean (6.14) and the solo mean (6.10) are statistically tied, so on score alone the panel added nothing. What it added was a bimodal split, Financial 8.2 against Execution 2.8, and the four errors above.

CartMySupply · panel 4.29 vs solo 5.00

ClassError the panel caughtCaught by
RevenueThe $2.7M headline is wrong by 7x to 10x against the proposal's own inputs, which compute to $269K. Stripe fees understated by roughly $11K per year. CAC absent entirely.Financial Integrity
CompetitiveTeacherLists already solves the identical problem, free, across 2 million lists. Target ships native School List Assist. No technical moat is claimed or demonstrable.Market Reality
ExecutionAmazon PA-API 5 removed Cart API support. Target has no self-serve multi-item cart API. Walmart requires separate catalog matching. The core mechanic of the product does not have a supported integration path at any of the three named retailers.Execution Feasibility
LegalCharitable solicitation registration required in 40+ states at $30K to $75K, plus COPPA exposure and FTC penalty risk. Compliance cost of $60K to $150K exceeds projected Year-1 revenue of $3K to $14K by an order of magnitude.Legal / Regulatory

Panel verdict: NO GO, unanimous, zero outliers, tightest spread of the three. Estimated rework 60+ hours. The build estimate of 116 hours was independently judged 4x to 10x too low.

Read the legal row again. A $60K to $150K registration obligation against $3K to $14K of projected revenue is not a scoring nuance. It is the difference between a business and a fine. A solo model reading the same document returned a 5.0 and did not raise it. That single finding is worth more than every point of score elevation the old thesis promised.

05Judge Pool v2.3: 11 Seats, 9 Vendors, Zero Double-Ups

The panel that produced the validation data ran at 8 of 11 seats because two seats hit a provider credit wall mid-run and one model family could not be dispatched at all. v2.3 is the spec written in response to that failure. Full detail lives in the judge pool specification v2.3.

The roster

BandSeatModelVendorScores
0Research AgentGrok 4.5xAINo
APrimary ReviewerClaude Opus 5AnthropicYes
ACross-Check ADeepSeek V4 FlashDeepSeekYes
ACross-Check BGemini Pro LatestGoogleYes
ACross-Check CDeepSeek V4 ProDeepSeekYes
ALegal / RegulatoryClaude Sonnet 5AnthropicYes
BFinancial IntegrityMiniMax-M3MiniMaxYes
BTeam / FounderClaude Fable 5AnthropicYes
BMarket RealityQwen3.7 PlusAlibabaYes
BExecution FeasibilityGPT-5.2 ProOpenAIYes
CSynthesis & Integrity GateKimi K2.6MoonshotNo

Nine distinct vendors across eleven seats. Nine distinct scoring models. Zero model double-ups: no single model occupies two scoring seats, which is the constraint that keeps correlated failure out of the panel mean. Maximum vendor concentration is Anthropic at 3 of 11 (27.3%), comfortably inside the 40% ceiling. DeepSeek holds 2 of 11 (18.2%). Every remaining vendor holds exactly one seat.

What changed in v2.3

Pre-flight health gate

Before any scoring begins, the orchestrator pings every rostered model with a 5-second probe and writes the result to a per-model health file. Any model returning HTTP 400, HTTP 429, or a no-healthy-deployments error is swapped for its pre-assigned failover before a single scoring call is spent. The 2026-08-12 run burned roughly 12 dispatches discovering dead models at runtime. That failure mode is now closed.

Pre-baked failover roster

Every seat carries a named failover from a different vendor, resolved at gate time rather than improvised mid-run. Failover selection preserves both the vendor-diversity ceiling and the no-double-up rule, so a degraded panel is still a valid panel rather than an accidentally correlated one.

Credit-wall resilience

The provider credit exhaustion that cost two seats mid-run is now detected at the gate and treated as an availability failure, not an error. Anthropic capacity has been restored and the Primary Reviewer seat runs Claude Opus 5 as specified.

Permanent exclusions

One frontier model family proved structurally incapable of running as a panel seat under our orchestration and is permanently excluded from the roster, not merely deprioritized. Excluded models cannot be selected as a failover target either.

Latency

v2.3 targets a critical path of roughly 113 seconds, against 218 seconds measured on the v2.1 architecture. The improvement comes from band parallelism: Band A and Band B seats execute concurrently rather than sequentially, and the Synthesis seat is the only stage that must wait for all scoring seats to return.

The ten scored dimensions

Every scoring seat rates the proposal 1 to 10 on the same ten dimensions, so panel spread is computed dimension by dimension and not only in aggregate:

Problem Clarity · Market Opportunity · Product Differentiation Revenue Model Viability · Go-to-Market Strategy · Competitive Moat Financial Projections · Team / Execution · Risk Mitigation · Legal / Compliance

The Synthesis and Integrity Gate

The Band C seat never scores. It reads all scoring output and performs a fixed checklist: verify score arithmetic, compute panel means and per-dimension spread, compute the delta against the solo baseline, flag every score more than 1.5 standard deviations from the panel mean with a written rationale, and confirm the verdict follows from panel evidence rather than from the Primary Reviewer alone. That gate is what turns eleven opinions into one auditable report.

06Worked Example: VentureBuilt v2, Where the Score Said Nothing

This is the clearest case in the validation set, because it is the one where score elevation delivered exactly zero and error detection delivered everything. Real scores from the 2026-08-12 run.

Panel mean · 8 judges
6.14/10

Median 6.45 · standard deviation 1.758 · spread 5.4 (min 2.8, max 8.2)

Solo baseline · Claude Opus 5
6.10/10

Spread 2.4 · delta +0.04 · statistically tied with the panel

Individual seat scores

SeatModelScoreSigma from meanFlag
Band B · Financial IntegrityMiniMax8.2+1.17Within 1.5 sigma
Band B · Market RealityGemini7.7+0.89Within 1.5 sigma
Band A · Primary ReviewerOpus 56.9+0.43Within 1.5 sigma
Band A · Legal / RegulatoryQwen6.5+0.21Within 1.5 sigma
Band A · Cross-Check CDeepSeek V4 Pro6.4+0.15Within 1.5 sigma
Band A · Cross-Check BGemini6.2+0.04Within 1.5 sigma
Band B · Team / FounderKimi4.4-0.99Within 1.5 sigma
Band B · Execution FeasibilityDeepSeek V4 Pro2.8-1.91OUTLIER
The average is a lie of composition. A 6.14 reads as a solid, fundable proposal with room to improve. The distribution says something completely different: the money is excellent (8.2) and the delivery plan is close to unworkable (2.8). Those are not two opinions about one thing. They are two accurate findings about two different things, and averaging them produces a number that describes neither.

What the outlier actually found

The 2.8 was not a grumpy model. The Integrity Gate challenged it at 1.91 sigma and it survived the challenge on evidence: a contractor budget broken 4x to 7x, an architecture that regressed from v1 with no schema and no API contract, and a founder workload of 37 engagements plus 950 hours alongside an MSP day job. The Team seat (4.4) and the Band A generalists (6.2 to 6.9) all acknowledged the same workload problem. They weighted it less severely. The Execution specialist is the only seat that forced it into the verdict.

Remove that one seat and the panel averages 6.6, which reads as a clean GO and ships a proposal with an 800-hour founder gap in it. The specialist seat cost a few cents of inference and changed the verdict.

Fix-It items, ranked by materiality

  1. CriticalExecution
    Fix the contractor budget or cut the scope. $1,500/mo buys 105 to 140 hours at market rates, not the volume the plan assumes. Raise to roughly $7,500/mo or reduce scope to fit the hours actually purchased.
  2. CriticalFinancial
    Restate Year 2 on a recognized-revenue basis. On run-rate the year looks healthy. On recognized revenue it loses roughly $11K to $18K. Present both.
  3. HighCompetitive
    Withdraw or qualify the uniqueness claim. LivePlan Plan Review and IdeaProof already ship in this space. Reposition on a defensible axis.
  4. HighTeam
    Name the contractor and the sourcing plan before Phase 2. A budget line with no named person is not a capacity plan.
  5. MediumGo-to-market
    Map the Year 1 to Year 2 GTM bridge. Eight net-new signups per month appear in the model with no acquisition mechanism behind them.
  6. MediumLegal
    Complete data protection and trademark clearance before Phase 0 to 1.

Panel verdict: CONDITIONAL GO. Estimated rework 15 to 20 hours. The solo baseline returned a 6.10 and none of the six conditions above.

07Pricing: Priced Per Error Found, Not Per Point Gained

Three tiers. The pricing logic follows the revised thesis directly: a panel run is worth what a caught error is worth, and a caught error is worth far more than a point of score.

Free

$0
  • 1 full review
  • Top 3 Fix-It items
  • Panel score and spread
  • Reduced panel size
Purpose: prove it on one document

Enterprise

$299/mo
  • Unlimited reviews
  • White-label branding
  • Multi-seat team workspaces
  • Configurable judge pool
  • Corpus isolation and data controls
  • Priority pipeline and support
Purpose: proposal teams running color reviews

What a review costs us, and why the panel is affordable

The validation run cost approximately $150 in inference for three full proposals across eight reporting seats, including retries against dead models before the health gate existed. That burn is the honest anchor for panel economics: a clean 11-seat run on one proposal, with the health gate preventing wasted dispatches, sits well inside single-digit dollars.

The reason a full 11-seat panel fits a $79 tier at five reviews per month is vendor mix. Only a minority of seats run premium frontier models. The specialist Band B seats run strong mid-tier models from five different vendors, which is where the error-detection value came from in validation. Panel diversity is cheaper than panel depth, and diversity is what caught the 14.

Why the value question is not the score question

Error classReal example from validationCost of missing it
Legal blockerCharitable solicitation registration in 40+ states$30K to $75K of registration, against $3K to $14K of projected revenue
Compliance totalFull first-year compliance load on the same proposal$60K to $150K, exceeding Year-1 revenue by roughly 10x
Execution gapContractor budget short by 4x to 7xRoughly 800 unbudgeted founder hours
Revenue arithmetic$2.7M headline against $269K computed from the document's own inputsCredibility with any investor who checks the math, which is all of them
Competitive blind spotA $4M-seed funded direct rival never named in the documentThe first question in the room, unanswered

A single caught item in the top two rows pays for a decade of the Pro tier. That is the entire pricing argument, and it does not depend on the panel producing a higher score, which it does not.

Positioned against the authoring category

ComparisonTheir priceVerdictTankMultiple
Pro vs Bidara Starter$499/mo$79/mo6.3x cheaper
Pro vs AutoRFP.ai Scale$899/mo$79/mo11.4x cheaper
Enterprise vs Bidara Starter$499/mo$299/mo1.7x cheaper
Enterprise vs AutoRFP.ai Scale$899/mo$299/mo3.0x cheaper

We are not a proposal team in a box. We are one high-value pass in the workflow. A buyer already spending $499 to $899 per month on an authoring tool should be able to add the error-detection layer without a second budget conversation. Pricing Pro at $79 makes VerdictTank an add-on decision rather than a platform decision.

08Competitive Landscape: Nobody Sells the Errors

Every AI-native player in this space is an authoring tool. They generate drafts. The nearest substitute for what we do is not a competitor product at all. It is a single frontier model and a prompt, and validation showed exactly what that substitute misses.

ProductCategoryPublished priceRelationship to VerdictTank
AutogenAIEnterprise authoringCustom, sales-led, no self-serveComplementary. We find the errors in what it writes.
CivioGov RFP authoringCustom, sales-ledComplementary. Downstream reviewer.
BidaraMid-market authoring$499/mo StarterComplementary. Transparent pricing, natural comparison anchor.
AutoRFP.aiResponse automation$899/mo ScaleComplementary. Reviews its drafts.
DeepRFPLean-team authoring$89/user/moComplementary. Lowest per-seat price in the category, natural partner.
A solo frontier modelDIY substituteAPI cost onlyThe real competitor. Measured: misses or underweights the material errors a panel catches.
VerdictTankPanel error detectionFree / $79 Pro / $299 EnterpriseThe only 11-seat, 9-vendor review panel with a published integrity gate

Competitor prices are vendors' own published rates as of July 2026. Tools without public pricing are shown as sales-led. Every named product was verified to exist and to occupy the authoring category.

Why the DIY substitute is the row that matters

Any buyer sophisticated enough to want proposal review can paste their document into a frontier model and ask for a critique. That is the honest competitive threat, and it is the one we tested against rather than around. The result is in section 3: the solo model returns a defensible, directionally correct score, and it returned none of the six VentureBuilt conditions, none of the CartMySupply compliance exposure, and none of the RFP Tank competitive reality.

1. Incumbents cannot sell honest criticism

Authoring tools sell the promise that they write your proposal. A brutal error list on the output that same tool just produced is a direct admission the generated draft is losing. It is structurally against their interest. We have no draft to defend. The verdict is the product.

2. Review is where the money is decided

Every serious bid already goes through a review gate, the color-team pass organizations run manually by pulling senior staff off billable work. That labor is expensive, slow, inconsistent between reviewers, and unavailable to the solo consultant. The demand is proven by the existence of the manual process.

3. Panel orchestration is a real moat

Nine vendors, health gating, pre-baked failover, no model double-ups, and an integrity gate that challenges its own outliers is not a prompt. It is an operations problem, and the 2026-08-12 run is the evidence of what it costs to learn.

4. An empty category sets its own price

A crowded category means the budget line exists and you fight for share. An empty review category means we define the line and set the reference price, while remaining complementary to every authoring tool in the table above.

The authoring tools write proposals. A solo model grades them smoothly.
VerdictTank tells you the fourteen things both of them got wrong.

09Deployment Options

Two supported deployment shapes. Both are managed by IT Pro Partner below the application layer.

Option A · ITPP-INFRA Shared

Runs on existing netcup RS 4000 infrastructure alongside IT Pro Partner operations. Same Wasabi S3 backup pipeline, same Caddy reverse proxy, same monitoring stack (Prometheus and Grafana). Zero new infrastructure cost. Suitable for launch through Series A.

  • netcup RS 4000 (app3), Docker Compose
  • Wasabi S3 daily backups plus 15-minute sync
  • Managed by the IT Pro Partner infrastructure team

Option B · Dedicated

Dedicated netcup or Hetzner instances with a dedicated S3 bucket. Full isolation from ITPP operational infrastructure. Recommended for post-Series A or enterprise white-label deployments requiring independent compliance scope.

  • Dedicated netcup RS or Hetzner CPX instances
  • Dedicated Wasabi S3 bucket, separate backup schedule
  • Managed by IT Pro Partner below the application layer

Shared responsibility: IT Pro Partner manages everything below the application layer (OS, container runtime, networking, backups, monitoring) under both options. The VerdictTank application and its model pipeline are the product team's responsibility.

Panel-specific operational requirement. Under either option the orchestrator must hold credentials for nine separate model vendors and must run the pre-flight health gate before every panel dispatch. Vendor credential rotation and per-vendor spend ceilings are application-layer concerns and sit with the product team, not with infrastructure.