Private beta — manually onboarded

Give every Claude request an owner. And every project a limit.

Put a control layer between your application and the Claude API. Issue project keys, set monthly token ceilings, and see what each request used—without storing its prompt in your usage logs.

Independent product. Not affiliated with, endorsed by or sponsored by Anthropic.

Request path Architecture diagram
Caller Your App server-to-server
Control layer GunTech Cloud bearer key check → monthly token reservation → rate window → forward
Upstream Claude API Anthropic Messages API
anthropic-version 2023-06-01

usage metadata (key reference, model, status, latency, token counters) is written back to the gateway ledger; the prompt body is not

Architecture diagram. A validated request travels from your application, through the gateway's key, policy and token checks, to the Anthropic Messages API. The gateway stores the request's token counters, not its prompt.
Problem

A prototype has one API key. Production has consequences.

The first version of a Claude feature fits in one file with one credential. The second version has environments, several services, and a customer who notices when something spends without a ceiling.

01

One credential serves every caller

A single upstream key is shared by staging, production, a script someone left running and the contractor who joined last month. When it leaks, there is nothing to revoke that isolates one caller without breaking all of them.

02

Usage you cannot attribute

A provider dashboard reports an account total. It does not say which application sent a request, which one retried in a loop overnight, or which one is safe to leave alone.

03

No ceiling before the call

Without a per-project limit, tokens are spent first and discovered later. The first alarm is an invoice, and by then the request has already been forwarded.

Capabilities

Three controls. Nothing decorative.

The gateway does three things: it issues project-scoped credentials, bounds token use per project, and records what each forwarded request used.

01

Project keys

Issue one gt_live_ key per project. Only a SHA-256 verifier is stored, the plaintext value is returned once at creation, and a key list never repeats it. Revoke a key and the next request from it returns 401 unauthorized.

revoke · rotate by issuing a new key

02

Token ceilings

Set a monthly ceiling per project. Before a request is forwarded, the gateway reserves the estimated input tokens plus the full max_tokens; if used + reserved + estimate would pass the ceiling it answers 429 budget_exceeded without calling Claude.

settle from reported usage · UTC month bucket

03

Request metadata

Every forwarded request writes one event: key reference, model, status, latency, input and output token counters, timestamp. The prompt body, the response body and credential material are not stored in that ledger.

inspect without storing prompts

max_tokens

16–4096, integer

messages

1–20, text only

system

≤ 12,000 characters

combined content

≤ 40,000 characters

monthly ceiling

10,000–100,000,000 tokens

rate limit

1–300 requests per minute

Technical proof

The request contract, as it ships

This is the request shape the gateway accepts today. model must match the model configured for this deployment; the value below is the model this deployment is configured with.

POST /v1/messages — request only
curl https://guntech.cloud/v1/messages \
  -H 'Authorization: Bearer gt_live_YOUR_PROJECT_KEY' \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "claude-sonnet-5-5",
    "max_tokens": 256,
    "messages": [{"role":"user","content":"Explain this error"}]
  }'

The block shows a request only. The reply is sent back exactly as Claude produced it, and no example response is printed anywhere on this page.

What comes back

The response body is the real Claude Messages payload, returned unchanged, including the usage counters the gateway settles against. No response body is reproduced on this page, because no response is fabricated here.

When something is wrong

If the deployment has no upstream credential, the gateway answers 503 not_configured instead of inventing an answer. Validation problems return 422, policy problems 429, and upstream failures 502 or 504 with a sanitised message.

The full route table, the error reference and every numeric constraint are in the API documentation.

Boundary

What it is. What it isn't. Yet.

Implemented means the matching source exists in this repository and can be verified locally. It does not mean the feature is deployed to guntech.cloud, and it does not mean a release gate has passed. This build is not certified for production.

Implemented in the repository

  • Public marketing website, documentation, terms and privacy drafts, robots and sitemap.
  • One-time bootstrap of a single administrator, with a PBKDF2 salted password hash and an HttpOnly session cookie.
  • Administrator project creation and listing, policy updates, and project API-key creation, listing and revocation.
  • Gateway POST /v1/messages for bounded non-streaming text against one configured model.
  • Conditional token reservation, monthly token usage tracking, fixed-window rate limits and response metadata events.
  • Expired session and reservation cleanup on a scheduled sweep.
  • Public pilot intake stored in the database, with Turnstile required for the production form.
  • Internal console showing real persisted data, not seeded customer or token figures.

Planned, not available

  • P1 Organizations, invitations, roles and per-organization tenancy; multi-administrator governance.
  • P1 Cost-estimation views from a versioned pricing catalog, and reconciliation against upstream billing.
  • P1 A full integration and regression suite, load tests, consistent request ids and retry guidance.
  • P2 Streaming, tool use, image requests, configurable model routing and idempotency support.
  • P2 Subscription billing, invoices, payments, self-serve signup, usage exports, notifications and webhooks.
  • P3 Additional providers, SSO/SAML, policy-as-code, enterprise features and formal compliance reviews.

Out of scope for v1: response caching, prompt storage or trace replay, fine-tuning, a guaranteed financial limit, token resale, official Anthropic affiliation, automatic customer signup and multi-model routing.

Release gates still open before an external private pilot: live domain and TLS, hosted database deployment, secrets and migrations, backup and restore proof, real upstream success and failure validation, attack-path testing, settlement-failure reconciliation, alerting and an incident runbook, legal identity and a monitored contact, and evidence of a working pilot tenant boundary.

Pilot intake

Request a pilot

Access is reviewed by a person. Send a work email and one concrete use case, and an operator will review it. Replies depend on the contact route described in the footer, which is still being verified. There is no self-serve signup, no public pricing page and no trial credential handed out automatically.

What happens next: an operator reads the request, checks whether the data you described can be handled in beta, and — if it fits — creates a separate project, sets its limits and delivers a project key through an approved channel.

This page loads no analytics script. Nothing you type here is sent anywhere except the intake endpoint of this deployment.

Stored only to review this request. The privacy statement lists each field and its purpose.

4–200 characters. One or two sentences. Do not paste customer data, prompts or credentials.

Manual review. No guaranteed response time.

FAQ

Questions a technical evaluator asks

Is the gateway live right now?

GunTech Cloud is in private beta, and this site does not publish an uptime figure or a latency number. Two public endpoints state the current condition instead: /api/health returns process health and reports upstream_configured separately, because a reachable process is not proof that an upstream credential exists; /api/public/config returns the model configured for this deployment and whether pilot intake is open.

Read either one yourself. A 200 from the gateway means Claude answered; if the credential is missing the gateway returns 503 rather than a fabricated completion.

Does the monthly token ceiling cap my spending?

No. Tokens are not money, and the ceiling is not a guaranteed monetary spending limit. The ceiling bounds the tokens the gateway observes, reserves and settles for requests that pass through it. Anthropic's own billing is separate and must be reconciled against your provider invoice. Usage from calls that never touch this gateway is invisible to it.

The arithmetic is explicit: before forwarding, the gateway reserves the estimated input tokens plus the full max_tokens; after the response it settles against the usage counters Claude reports. A request that would push used + reserved + estimate past the ceiling is rejected with 429 budget_exceeded before anything is forwarded.

Do you store my prompts?

Not in the usage ledger. Each forwarded request writes one event holding a key reference, the model, the HTTP status, the outcome, the latency in milliseconds, input and output token counters, and a timestamp. The prompt body, the response body and credential material are not stored in that ledger. Infrastructure logs are single-line JSON records carrying an event name, a correlation id, status, outcome and latency.

If you need prompt storage or trace replay, this gateway is the wrong tool for that job.

Can I stream responses, use tools, or choose a model per request?

Not yet. One model is configured per deployment, and the endpoint accepts non-streaming plain-text requests only. stream, tools, tool_choice, image content blocks and any unrecognised top-level parameter are rejected with 422 validation_error before anything reaches Anthropic.

Model routing, batches and idempotency keys are planned work, not available behaviour.

Is this an official Anthropic product?

No. GunTech Cloud is an independent product. It is not affiliated with, endorsed by or sponsored by Anthropic. Claude is a product of Anthropic; this gateway forwards validated requests to the Anthropic Messages API using a credential held by the operator, and Anthropic's terms govern that upstream usage.