Private Generative AI: modern LLMs on your internal data, without it leaving your controlled cloud
Approve a pilot to give employees modern large language models — Anthropic's Claude, hosted on AWS — that process the organization's most sensitive internal data entirely inside our own AWS environment in Canada, with the data never reaching Anthropic's or OpenAI's cloud.
Sensitive internal data is currently off-limits to AI
Today, employees are effectively blocked from using modern generative AI on real work, because sending internal information — call transcripts, financial records, customer and employee personal data, source code and intellectual property, contracts, and strategic plans — to a public AI service was never an acceptable privacy risk. So the organization has either banned these tools or restricted them to trivial, non-sensitive tasks. The capability gap keeps widening while the data that would make AI genuinely useful stays off-limits.
This proposal closes that gap. It stands up a private deployment of the native Claude desktop application in which the model runs on AWS Bedrock inside our own AWS account, in Montreal. Employees get a first-class AI assistant that can summarize call transcripts, analyze company data assets, draft and review documents, and reason over internal systems — while the data those tasks depend on stays within our controlled cloud environment in Canada. This is the first time the organization can safely point frontier AI at its own sensitive data.
A precise claim, stated honestly: the data does not reach Anthropic or OpenAI, does not traverse the public internet, and does not leave Canada. It is processed by AWS Bedrock, a service running inside our own AWS account under our own encryption keys. AWS operates that infrastructure under a shared-responsibility model, the same way it already hosts regulated workloads for financial, healthcare, government, and legal organizations. AWS Bedrock does not use our prompts to train models and does not share them with the model provider.
The request never leaves our AWS account in Canada
Inference travels over a private network path into our own AWS account and returns — the AI vendors' clouds and the public internet are never a hop on that path.
Anthropic's cloud(model provider)OpenAI's cloudThe public internetAny region outside Canada
What stays inside our control
- Prompts, call transcripts, company data assets (financials, PII, IP, contracts)
- Model inputs and outputs
- Encryption keys (customer-managed, our KMS)
- Full audit trail (our logs, our region)
- The AWS account and network the model runs in
What never leaves / never sees the data
Anthropic's cloud(model provider)OpenAI's cloudThe public internetAny AWS region outside CanadaAny human at Anthropic or OpenAI
Five building blocks, all inside ca-central-1
The Technical Implementation Blueprint specifies each of these in engineering detail.
Managed desktop client
The native Claude desktop app, centrally configured to run in "third-party inference" mode against our AWS account and locked down so users cannot repoint it or sign in to a public account.
AWS Bedrock (model hosting)
Runs the Claude model inside our account in Montreal. Frontier-class models, with the ability to adopt newer models as AWS releases them — no in-house model training or hosting required.
Private network path (AWS PrivateLink)
Inference traffic travels over AWS's private backbone via a VPC endpoint, never the public internet.
Customer-managed encryption (AWS KMS)
Our own keys wrap the data and the audit trail; we hold the authority to grant or revoke access.
Audit & residency enforcement
Every request is logged to our own store in Canada, and an account-wide policy guardrail refuses any attempt to run the model outside ca-central-1.
Provable residency, and we hold the keys
- Data residency is provable, not promised. Processing and logs are pinned to Montreal and enforced by policy, so we can demonstrate to an auditor that a request could not have run outside Canada.
- Zero exposure to the AI vendors' clouds. Anthropic and OpenAI are removed from the data path entirely, and our data is never used to train their models.
- We hold the keys. Customer-managed KMS means the organization — not a vendor — controls who can decrypt prompts, outputs, and logs.
- Full attribution and audit. Per-user sign-in through our corporate identity provider means every AI request is attributable to a named employee in our own logs.
- Unlocks internal data for AI. For the first time, sensitive internal content can be used with modern LLMs, turning a previously off-limits data estate into a productivity asset.
Every material risk has a concrete control
No architecture is risk-free. The material risks are known and each has a concrete control; the Technical Implementation Blueprint details them in full.
| Risk | How it could fail a security check | Control |
|---|---|---|
| Cross-region data leakage | AWS "cross-region" inference can serve a request outside Canada | Mandate in-region invocation; account-wide policy (SCP) hard-denies any Bedrock call outside ca-central-1 |
| End-user tampering | A user edits local settings to repoint the app or reach a public endpoint | Configuration deployed as tamper-resistant machine policy; in-app settings become read-only; public sign-in path disabled |
| Shadow AI / bypass | A user pastes data into public ChatGPT or Claude in a browser | Sanctioned tool made genuinely capable + corporate egress controls; adoption is the real mitigation |
| Data exfiltration via plugins/tools | User-added extensions or shell/computer-use tools move data out | High-risk tools and user-added plugins disabled; sandbox egress allow-listed to our endpoints only |
| Endpoint compromise | This design protects the inference path, not the laptop itself | Existing device security (encryption, EDR, MDM) remains required and is a prerequisite, not replaced |
| Vendor launch dependency | The app fetches its launch bundle from an Anthropic host (downloads.claude.ai) — app code only, no data | Documented for auditors; the offline installer removes this host entirely |
Cost scales with usage, not infrastructure
For 100 users over 12 months, the modeled cost is ≈ $43,320 USD/year. Roughly 97% of that is variable, usage-based model consumption (billed per unit of text processed); the fixed cloud infrastructure — private networking, encryption, logging, identity — is only ≈ $110/month (≈ $1,320/year).
Practically, cost scales with actual use, and the biggest savings lever is usage efficiency — prompt caching can cut input costs by up to ~90% — not infrastructure trimming. Treat the figure as a budgeting envelope with a usage-driven variance band, not a fixed subscription. The full model and assumptions are in the Technical Implementation Blueprint.
On accuracy: these are the organization's own planning estimates for modeling — not figures quoted by AWS, Anthropic, or OpenAI. On-demand inference is consumption-billed and variable; confirm live rates and ca-central-1 model availability in the AWS console before committing. See References & Sources → Note on figures.
OpenAI models on the same architecture
The identical private architecture can run OpenAI's GPT-5.6 models on AWS Bedrock instead of Claude, with the same Canadian residency, private networking, and encryption controls — modeled at ≈ $35,520 USD/year, roughly 18% lower on model-consumption pricing.
The trade-off: unlike Claude, OpenAI has no native, centrally-managed desktop application for this deployment, so the polished locked-down desktop experience, MDM controls, and governance tooling described here would have to be sourced or built separately — offsetting part of the token saving and adding schedule risk. Recommendation: settle the choice with a scored pilot, not on price alone.
Proceed to a funded 100-user pilot
Proceed to a funded 100-user pilot on the Claude-on-AWS architecture in ca-central-1. It gives employees modern AI on real internal data for the first time, keeps that data inside our controlled cloud environment in Canada with provable residency, and does so at a predominantly usage-based cost with a fast path to production. Confirm final pricing and model availability in the live AWS console, and validate the residency controls against our compliance requirements, before committing to full rollout.
The full detail, in three companion pages
Technical Implementation Blueprint
Architecture, AWS controls, client & MDM configuration, the cost model, and the full risk analysis.
Alternative Platform: OpenAI Analysis
OpenAI GPT-5.6 on the same perimeter — the side-by-side comparison and where the trade-offs land.
References & Sources
Every vendor source behind the architecture, plus the important note on how the figures were derived.