The multi-tenant internal staff knowledge-base chat platform - a private, cited AI for your employee handbook, billing rules, and policies.
Wall-O is a multi-tenant internal staff knowledge-base chat platform. Every knowledge domain inside a business (employee handbook, billing, IT help, policies, procedures) becomes its own scoped AI chat channel, backed by that business's real documents. Staff ask questions in a chat interface and get grounded, cited answers drawn from the company's own files - never hallucinated, never leaking across domains.
The core problem is universal and business-agnostic: employees waste a documented, measurable fraction of their week hunting for information that already exists somewhere in the organization. McKinsey and Gartner studies put that fraction in the double digits of weekly hours. Enterprise solutions (Glean, Moveworks, Coveo) solve this for large companies at enterprise prices and enterprise procurement complexity. The SMB and vertical-niche middle - dental practices, law firms, restaurants, retail groups, logistics firms - is priced out and under-served.
Wall-O's answer is a hard invariant: one channel equals one knowledge domain, one attached AI agent, and one scoped knowledge source. This makes answers provably grounded (every reply carries citations to the source document) and provably isolated (a billing question can never surface a patient file or a neighboring tenant's data). The product is self-hosted, white-label, and business-agnostic by construction - the same core serves a 12-person dental practice and a 200-person law firm.
Financially, Wall-O is a high-margin software product running on infrastructure IT Pro Partner already owns: ~$68K build cost, ~$0.50 per 1,000 queries marginal cost, ~95-98% gross margin, ~$1,000 CAC against ~$14K LTV, and full build-cost payback in ~10-13 months (realistic). The v1 is a grounded Q&A assistant. The v2 roadmap adds risk-tiered "do" capabilities (reporting, content generation, integrations) that deepen the moat without ever touching patient/PHI scope.
This proposal defines v1 and v2 concretely, prices it against the market, and makes the fact-based case for building it. It is submitted for critical review.
Wall-O is the AI coworker that reads your business's actual documents and answers your team's questions with citations - for the price of a streaming subscription, on infrastructure you control, scoped so nothing leaks. It is "a private, cited ChatGPT for your employee handbook, billing rules, and policies" - one tightly-scoped assistant per knowledge domain, white-labeled to your brand.
Every business with more than a handful of employees has the same problem, regardless of industry:
This is not an industry-specific problem. A dental practice has the same shape of problem as a law firm, a restaurant group, a retail chain, a logistics company, or a real estate office. The documents differ; the pain is identical.
| Current method | Time cost | Pain level | Why it fails |
|---|---|---|---|
| Ask a manager / senior coworker | High (interrupts them every time) | High | Doesn't scale; same questions repeat |
| Search SharePoint / Google Drive manually | Medium (10-20 min per hunt) | Medium | Fragmented, no ranking, poor recall |
| Printed binders / shared docs | Medium | Medium | Goes stale immediately |
| Generic LLM (public ChatGPT) | Low | High (risk) | Hallucinates, no access to internal docs, leaks data |
| Enterprise knowledge AI (Glean, Moveworks, Coveo) | Low | Low (but expensive) | $30+/seat/mo, enterprise procurement, overkill for SMB |
The gap: there is no affordable, self-hosted, white-label answer for the SMB and vertical-niche middle. Enterprise tools solve the problem but are priced and packaged for enterprises. Public LLMs are cheap but ungrounded and unsafe for internal documents. Wall-O sits in that gap.
The buyer is the owner or office manager of an SMB (10 to 200 employees) who:
The beachhead vertical is orthodontics/dental (Wall Orthodontics), but the product model is explicitly business-agnostic - the same channel primitive serves any document-driven business.
| Need | Public LLM | Enterprise AI | Wall-O |
|---|---|---|---|
| Grounded, cited answers from my docs | No | Yes | Yes |
| Affordable for SMB | Yes (but unsafe) | No | Yes |
| Self-hosted / data control | No | Partial | Yes |
| White-label to my brand | No | Partial | Yes |
| Channel-scoped (no cross-domain leak) | No | Partial | Yes (hard invariant) |
| No PHI / patient scope creep | No guarantee | No guarantee | Enforced at ingestion |
Methodology: per-seat SaaS anchored on knowledge-worker count ("every employee could use an internal Q&A assistant"), blended at $10/user/month - well below the enterprise incumbents that charge $30-75 with minimums.
US-only cross-check (tighter, verifiable): 100M US knowledge workers x 45.9% = ~46M SMB workers x $120/yr = **$5.5B US SMB SAM**. 1% capture = $55M/yr; 5% = $275M/yr. Even a single vertical slice (e.g. ~1M US small professional-services firms) is a meaningful SOM.
Sources: ~1B global knowledge workers (Schroders); ~100M US knowledge workers (Upwork/BLS/Eurostat); SMBs employ 45.9% of US private-sector workers (SBA 2024).
These are the proof points that the category is real, funded, and acquirers pay billions for it:
| Company | Signal | Number |
|---|---|---|
| Glean | Raised $765M, valued **$7.2B** (Series F, Jun 2025) | Passed $100M ARR 2025; claims up to **110 hours saved/user/year** |
| Moveworks | Acquired by ServiceNow for **$2.85B** (Mar 2025) | Customers HP, Unilever, Toyota, Marriott; **70,000 hours reclaimed** at one automaker; 75,000 hours at a biopharma |
| Microsoft 365 Copilot | **20M+ paid seats** (of 450M M365 seats) | 75% of knowledge workers already use gen AI at work (MSFT Work Trend Index) |
| Notion AI | **$500M annualized revenue** (Sept 2025) | AI add-on attach rate grew 10-20% to 50%+ in one year |
On the "every business has this problem" claim, the independent studies converge regardless of industry or size:
These cluster around 20-30% of the workday lost to search - a universal, business-agnostic cost, not an enterprise-only one. That is the budget Wall-O attacks.
Full 10-competitor analysis is in the research appendix. The key fact: **the $300-800/month self-hosted band for 10-30 seat SMBs is structurally empty.**
| Competitor | Pricing | Target | Wall-O's gap |
|---|---|---|---|
| Glean | $50-75/seat, 100-seat min (~$60K/yr) | Enterprise | Self-host + white-label + vertical templates + SMB price |
| Moveworks/ServiceNow | $100K+/yr | Enterprise IT/HR | Not ticket-automation; general internal knowledge |
| Guru | $10-25/seat, 10-seat min | Mid-market | No self-host, no white-label, no channel scoping |
| Notion AI | ~$10/seat add-on | General workspace | Q&A only over Notion content; no M365 connector |
| M365 Copilot | $21-30/seat + M365 base | M365 tenants | Microsoft-cloud lock-in; no self-host/white-label/scoping |
| Coveo | Enterprise QPM | Fortune 1000 | No SMB tier |
| Dust | $30-150/seat credits | AI-operator teams | DIY agent-builder, not turnkey |
| Open WebUI / Onyx / AnythingLLM | Free self-host | Developers | Raw RAG kits - no multi-tenant isolation, no M365 connector, no vertical layer |
Two strategic notes from the research:
Every Wall-O channel is exactly three things bound together, and this is a hard invariant:
A channel has exactly one agent and exactly one scope. A channel cannot span two domains, and an agent cannot serve two channels. If a business needs a second domain, it creates a second channel, a second agent, and a second scope. This invariant is what makes answers grounded (only one scope's documents feed the answer) and isolated (no cross-domain or cross-tenant leak), and it is enforced in the schema with unique constraints running in both directions.
| Layer | Responsibility | Owns |
|---|---|---|
| Rocket.Chat | Chat transport only | One workspace per tenant (MIT core, EE stripped). Rooms, users, messages. Zero intelligence. |
| Orchestrator | All intelligence | Multi-tenant FastAPI service. Tenancy, agents, kb_scope, M365 connector, retrieval, LLM, posting. |
| Postgres + pgvector | State and vectors | Tenants, channels, agents, scopes, documents, chunks, messages. |
| admin-ai (LiteLLM) | LLM | DeepSeek V4 Pro primary, configured fallback chain. |
| Wasabi S3 | Object storage | M365 sync staging, backups, agent assets, audit exports. |
The intelligence never lives in the chat layer. Rocket.Chat is a dumb transport; the orchestrator owns everything that thinks. This split is what makes multi-tenancy clean and white-labeling a server-side concern rather than a per-tenant code fork.
v1 is the answer engine, nothing more. A staff member @mentions the channel's agent (or DMs it), and gets back a grounded, cited answer drawn from that channel's scoped documents.
**The v1 message flow (6 steps):**
**v1 hard behaviors (SETTLED):**
**What is already built vs. to build (v1):**
| Component | Status | Effort |
|---|---|---|
| Product model + data architecture | SETTLED (3 design docs committed) | Done |
| Data model (tenants, channels, agents, kb_scopes, documents, chunks, messages, users) | Spec'd, not built | Build |
| Orchestrator (FastAPI, tenancy, retrieval, LLM client) | Not built | Build |
| Postgres + pgvector schema | Spec'd, not built | Build |
| Rocket.Chat workspace + bot integration | Phase 0 checklist spec'd, not built | Build |
| M365 connector (Sites.Selected read) | Spec'd, not built | Build |
| White-label server rebrand (fossify FOSS build) | Phase 1 gate, not started | Build |
| Auth (Hexclave/Stack Auth, Entra OIDC for O365) | Partially existing (Hexclave on app3) | Integrate |
**Phase 0 proof slice:** stand up one Rocket.Chat workspace for Wall Orthodontics, register a bot, create a channel, attach an agent, index one document, and close the loop on a single grounded answer. The 21-step Phase 0 checklist is spec'd and ready to execute.
v2 adds risk-tiered "do" capabilities on top of the v1 answer engine. These are deliberately ordered by write-risk, and none of them ever unlock patient/PHI or autonomous destructive action.
| Capability | Tier | Write risk | Auto-approve? | Value |
|---|---|---|---|---|
| Reporting (structured extraction + SQL aggregation + chart render) | Reporting | Zero new write risk | Yes | Highest value, zero risk - recommended first v2 ship |
| Content generation (drafts, summaries, boilerplate) | Content generation | Drafts only | Yes | Saves drafting time |
| Housekeeping (move/rename/delete-to-recycle) | Housekeeping | Destructive | No - propose/approve/audit | Keeps knowledge fresh |
| Integrations (Graph delegated sendMail/calendar/webhooks) | Integrations | External side effects | No - approval | Connects Wall-O to workflows |
| Design/brand (template render + image gen) | Design/brand | Drafts only | Yes (template-driven only) | Flyers, branded assets |
**Write rails (SETTLED):** destructive actions require propose -> approve -> narrow audited write with before/after and undo path. Harmless writes auto-approve by risk tier. Per-agent tool scoping preserves one-agent-one-domain. Template-driven rendering is required for text-accurate flyers; raw image generation is unreliable for text and is not shipped for that use.
**Permanently out of scope (never unlocks):** patient records, PHI, HIPAA scope, clinical/x-ray image analysis (separate FDA-regulated product), and autonomous destructive actions. These are hard rails, not feature gaps.
The v1 and v2 product model references no industry. "Tenant" is any business; "channel" is any knowledge domain; "documents" are any document type. The M365 connector reads SharePoint/OneDrive libraries generically. The only orthodontics-specific artifacts are the pilot client's name and branding. A law firm, restaurant group, or logistics company is onboarded by the same provisioning path with different documents, different channel names, and different branding - zero code change.
Premium positioning, not race-to-bottom. Wall-O is "the most accurate, best-scoped internal knowledge assistant", not the cheapest. Price for value delivered (time saved, errors avoided), not competitor undercutting. The product is self-hosted software with high gross margin, so price is set by willingness-to-pay in the segment, not by our marginal cost.
| Tier | Price | Includes |
|---|---|---|
| Starter | $199/tenant/mo (up to 10 users, +$15/user beyond) | 1 M365 connector (SharePoint + OneDrive), Rocket.Chat, 50K docs, 1 channel scope set |
| Business | $499/tenant/mo (up to 25 users, +$25/user beyond) | + Teams connectors, unlimited channel scoping, SSO (OIDC/SAML), 250K docs, SLA support |
| Enterprise | $2,000/tenant/mo (annual contract) | Dedicated instance, unlimited seats, custom connectors, SCIM, DPA, custom retention, white-glove onboarding, 99.9% SLA |
Blended ARPU at a 60/30/10 tier mix = $469/mo (modeled at $450 for safety). Every tier includes the core grounded Q&A engine, the channel-scoped isolation invariant, citations, audit logs, and the no-PHI guardrail - higher tiers unlock more v2 "do" capabilities and white-glove onboarding, not better core answer quality.
The one-channel-one-agent-one-scope invariant is not a feature a competitor can bolt on after the fact; it is the schema. It makes every answer provably grounded and provably isolated, and it is what lets Wall-O serve regulated-adjacent verticals (dental, legal) where data leakage is a dealbreaker. Enterprise tools do broad search over everything; Wall-O does scoped answers over exactly one domain.
Wall-O pairs SMB affordability with the isolation discipline usually found only in enterprise platforms. That combination is the open gap in the market.
Data never leaves the customer's estate (or ITPP's controlled netcup estate). Chat, documents, and vectors stay self-hosted; only the LLM call routes through the internal LiteLLM proxy, and by scope it never carries PHI. White-labeling is a server-side rebrand (fossify FOSS build), so every tenant feels like their own branded product, not a resold third-party tool.
IT Pro Partner already runs the infrastructure and has the MSP relationship model. Wall-O is not a cold-start consumer product; it slots into an existing managed-services motion where the first customer (Wall Orthodontics) is already warm.
The risk-tiered "do" capabilities (reporting first, then integrations) deepen switching cost once a tenant's staff builds habits and workflows on the assistant. Reporting is the recommended first v2 ship because it is highest value with zero new write risk.
Prove the loop with one real client in one vertical. Phase 0 closes a single grounded answer; Phase 1 delivers a white-labeled workspace with 3-5 channels (Employee Resources, Billing, IT Help, and a marketing/design channel as a tool-belt extension). Collect concrete ROI (questions answered, time saved, onboarding speed) for the case study.
The orthodontics win becomes a repeatable template for adjacent verticals: general dental, then any document-driven SMB (law firms, accounting, real estate, restaurants, logistics). Each vertical gets a tailored channel template (e.g. "billing" channel for a dental practice vs. "matter intake" for a law firm) but zero core code change - the business-agnostic model pays off here.
Sell through IT Pro Partner's managed-services relationships and peer referral (practice-to-practice, firm-to-firm). The pitch is "the AI that reads your actual documents", demonstrated with a live tenant, not a slide deck. White-label means each MSP or franchise group can offer it under their own brand.
| Risk | Likelihood | Impact | Mitigation |
|---|---|---|---|
| LLM answer quality (hallucination) | Medium | High | Never-fabricate rail, citation requirement, below-threshold "I could not find an answer", model fallback chain |
| Data leakage across tenants/channels | Low | Critical | Hard invariant, tenant_id on every row, RLS, vector namespace isolation, HMAC webhook auth |
| PHI accidentally ingested | Low | Critical | Allowlisted libraries only, PHI-marker pre-ingest filter, block-and-log |
| M365 Graph API rate limits / connector fragility | Medium | Medium | Delta sync, exponential backoff, Graph search fallback |
| Rocket.Chat EE licensing in resale | Low | High | fossify FOSS-only build, never ship stock EE image |
| Native app store cost (Phase 2) | Medium | Low | PWA first, native only after PWA validated |
| Slow build (scope creep into "do" features) | Medium | Medium | v1 is answer-only; v2 capabilities are risk-tiered and gated |
| Competition (enterprise tools move downmarket) | Medium | Medium | SMB price point + isolation + white-label + MSP distribution |
| Churn if ROI not demonstrated | Medium | High | Measure time-saved from day one, report it to the buyer monthly |
Hallucination. If Wall-O ever gives staff a wrong, confidently-stated answer about policy or billing, trust collapses and the product is done. The never-fabricate rail, mandatory citations, and below-threshold "I could not find an answer" behavior are not features - they are the product. This is why the answer engine ships before any "do" capability, and why the v2 write features are risk-tiered with approval gates.
Assumption: $100/hr blended engineering rate (loaded senior full-stack or mid contractor), no new hardware (runs on existing netcup estate).
| Component | Hours | Note |
|---|---|---|
| M365 Graph connector (ACL mapping, delta sync, webhooks) | 145 | Hardest single line |
| Ingestion pipeline (parse, chunk, embed, pgvector) | 70 | |
| RAG retrieval + grounded generation + citations | 90 | |
| Multi-tenancy (isolation, config, vector namespaces) | 50 | |
| Rocket.Chat integration | 70 | |
| Auth + SSO (Stack Auth / Hexclave) | 70 | |
| Admin UI + billing + onboarding | 90 | |
| Testing, security, observability, deploy, docs | 90 | |
| **Total** | **675** |
**Build cost = 675 hrs x $100 = ~$68K** (range $45K-$100K). Ongoing engineering ~$3,500/mo post-launch.
DeepSeek V4 Flash via LiteLLM ($0.14/M in, $0.28/M out). Per answer: ~2,000 input + ~500 output tokens.
Sensitivity: even at V4 Pro rates ($0.435/$0.87) it is $1.31/1K queries; a 5x price rise is ~$2.10/1K queries. All negligible. Per-tenant COGS is ~$3-30/mo depending on tier. **Gross margin ~95-98%.**
| Metric | Value |
|---|---|
| Blended ARPU | ~$450/mo (tier-mix math gives $469) |
| Blended COGS/tenant | ~$7/mo |
| Gross margin | ~95-98% |
| Churn | 3%/mo base (5% conservative) |
| LTV | ~$14K (range $8.5K-$21K) |
| CAC | ~$1,000 (owned MSP distribution is the moat) |
| LTV:CAC | ~14:1 (healthy SaaS is >3:1) |
| Payback per tenant | ~2.3 months |
| Scenario | Tenants (month 12) | MRR (month 12) | 12-month revenue |
|---|---|---|---|
| Conservative | 22 | $9,900 | ~$52.7K |
| Realistic | 48 | $21,600 | ~$107.6K |
| Aggressive | 90 | $40,500 | ~$198.9K |
**Yes - build it.** The math is unambiguous because three things compound:
Full payback in ~10-13 months realistic, profitable even in the conservative case within ~2 years. The honest caveats: the M365 connector is the riskiest engineering line (fund and pilot it first), DeepSeek may raise prices (LiteLLM abstraction absorbs it), and churn/adoption are unproven in a new category (activation, queries-per-active-user, is the KPI to watch).
The build requires, and the design docs already specify:
| Infrastructure | Status | Detail |
|---|---|---|
| netcup host for Wall-O | To provision | Docker host, Caddy TLS edge, one Rocket.Chat workspace + MongoDB per tenant |
| Orchestrator (FastAPI) | To build | Multi-tenant, owns tenancy/agents/scopes/retrieval/LLM |
| Postgres + pgvector | To build | Tenancy tables, chunks + HNSW vector index |
| admin-ai (LiteLLM) | Existing | DeepSeek V4 Pro primary + fallback chain |
| Wasabi S3 | Existing | Sync staging, backups, assets, audit exports |
| Vaultwarden | Existing | All secrets by reference, never inline |
| Hexclave/Stack Auth (app3) | Existing | Customer-facing SaaS auth; Entra OIDC for O365 pilot |
| White-label FOSS build | To build (Phase 1 gate) | fossify script + CI image build, never ship stock EE image |
**Option A - ITPP-INFRA Shared:** runs on the existing netcup estate alongside ITPP operations, backed by the same Wasabi S3 backup pipeline. Lowest cost, fastest start, ITPP manages everything below the app layer.
**Option B - Dedicated:** dedicated netcup/Hetzner instances with a dedicated S3 bucket, managed by ITPP, for tenants that demand physical separation or for the multi-tenant production fleet at scale.
Shared responsibility is explicit: ITPP manages everything below the app layer (servers, Docker, TLS, Postgres, backups, monitoring); the customer owns their documents and how they use the assistant.