title: AI Usage Tracking source_url: /developer-api/v1/ai-usage summary: Add a Ramp broadcast destination to your AI platform so customers can bring their Ramp API key and receive their own model usage, token, cost, and attribution data in Ramp. content: Add a Ramp broadcast destination to your AI platform so customers can bring their Ramp API key and receive their own model usage, token, cost, and attribution data in Ramp. An inbound broadcast integration for your customers: Your product adds a Ramp destination where each customer pastes or connects their own Ramp AI usage API key. Your platform securely stores that key for the customer and sends that customer's usage events to POST /developer/v1/ai-usage/unified. Each event includes model, provider, token usage, pricing context, optional reported cost, optional latency, and customer attribution metadata. Ramp ingests the events into the customer's AI spend analytics without requiring them to poll invoices, scrape dashboards, or maintain per-provider ETL. Your company operates an AI platform, model provider, proxy, or telemetry service that meters customer usage. You can add a customer-facing Ramp destination setting and securely store a Ramp API key per customer. Your customer has a Ramp production account and can generate a Ramp AI usage API key scoped to ai_usage:write. Your usage pipeline can send an HTTPS POST request after each completed model call or as an asynchronous batch flush. Implementation In your product, add a Ramp destination setup flow. Tell customers to open Ramp, go to Settings > Integrations, find your integration, click Connect, and generate an AI usage API key. The customer copies that key into your destination settings. Store the key as a customer-scoped secret and use it only for that customer's usage events. If the customer rotates or revokes the key in Ramp, prompt them to reconnect the destination. Send unified usage events to: Authenticate each request with the customer's Ramp API key: Each request is a batch envelope. Use one event per model call, and batch multiple events when your platform flushes usage asynchronously. Do not mix events from different Ramp customers in the same request. Ramp returns 204 No Content after accepting the batch for asynchronous ingestion. Top-level field Type Required Description schema_version string Yes Unified schema version. Send 1.0; unknown versions are rejected. events array AI usage events. Each item represents one model call. Identity fields make the event deduplicable and explain where it came from. Field event_id Stable caller-provided id for this model call. Reuse the same id and unchanged event data on retries so Ramp can deduplicate delivery. source Your emitting system or platform name. occurred_at datetime When the model call happened. Use an ISO 8601 timestamp with timezone; naive datetimes are rejected. provider Model provider name. model Provider model name used for the request. gateway_key_id No Identifier of the LLM Gateway key associated with this event. gateway_key_alias Gateway key alias captured when this event occurred. Use usage for token counts and other billable meters. All token fields are required and must be non-negative integers. Send 0 when a token bucket is known to be zero. usage.input_tokens integer Prompt/input tokens billed or counted for the request. usage.output_tokens Completion/output tokens billed or counted for the request. usage.cache_read_input_tokens Input tokens served from cache. usage.cache_write_input_tokens Input tokens written into cache. usage.reasoning_output_tokens Reasoning tokens produced by models that expose a separate reasoning-token counter. usage.meters Additional provider-specific usage meters when a provider bills on dimensions beyond tokens. Send [] when there are no extra meters. Use pricing_context for the inputs Ramp needs to price or explain pricing differences. service_tier, fast_mode, session_created_at, and long_context are required. Use reported_cost when your platform already calculated a cost for the event; omit reported_cost if you do not have a caller-computed cost. pricing_context.service_tier Provider service tier or priority lane, such as standard, priority, or batch. pricing_context.fast_mode boolean Whether the request used a low-latency or premium-speed path. pricing_context.session_created_at Session creation time when session age affects pricing or attribution. Use an ISO 8601 timestamp with timezone; naive datetimes are rejected. pricing_context.long_context Whether the request used a long-context pricing mode. pricing_context.inference_geo Optional inference region or geography when regional pricing applies. Defaults to an empty string. reported_cost.amount decimal Required when reported_cost is provided. Non-negative cost you calculated for the customer event, up to 20 digits with 10 decimal places. Send as a string to preserve precision. reported_cost.currency Required when reported_cost is provided. Currency for the reported cost. Use USD. reported_cost.provenance Required when reported_cost is provided. How the cost was calculated, such as provider, platform, or estimated. reported_cost.estimated Required when reported_cost is provided. Whether the reported cost is estimated rather than final. Pass stable attribution with each model call so Ramp can break down spend by the dimensions your customer uses. All attribution fields are required; use opaque identifiers instead of emails or names unless your customer explicitly asks you to send richer identifiers. provider_latency_seconds number Provider-reported model response latency, if available. attribution.session_id Conversation, workflow, or job session id. attribution.turn_id Turn, step, or request id within the session. attribution.user_id Stable customer-side end-user or service-user identifier. attribution.tags object String-to-string metadata supplied or derived for that customer's team, project, environment, cost center, feature, use case, or department. Prefer metadata values that are stable and safe to aggregate. Avoid sending secrets, API keys, or unnecessary prompt and completion content. Extra fields that are not part of the schema are ignored. Use retries. Retry temporary failures, such as 429, 5xx, and network timeouts, with exponential backoff. Fix invalid requests or authentication failures before retrying. Reuse the same event_id and original event payload on every retry. Changing usage, cost, metadata, or other event data can create a separate spend record. Store customer keys securely. Treat each Ramp API key like a customer secret. Encrypt it at rest, restrict access, and support rotation or deletion. Scope by customer. Send each customer's events with that customer's Ramp API key. Never batch events for different Ramp customers under one key. Tag every workload. Untagged usage still appears in aggregate spend, but team, project, and cost-center views depend on attribution tags. Avoid prompt content. The unified schema is designed for usage, pricing, and attribution metadata; do not send prompts or completions. Validate with the customer. Send a test model call through your platform and have the customer confirm that usage appears in their Ramp AI Spend dashboard. AI Agents - connect agents to Ramp for actions and auditability after tracking their model spend. AI Index - market-level AI adoption metrics from Ramp's aggregate spend data. Ramp Rate - software adoption and growth metrics across Ramp's network. POST /developer/v1/ai-usage/unified - ingest unified AI usage event batches. Authorization - authenticate API calls and understand OAuth scopes. Errors - handle non-2xx responses and retry safely. Rate limits and timeouts - design callback retry behavior around API limits.