Sample trace format
Send a sanitized request trace in the format below, plus a short fleet note. We’ll arrange a suitable way to transfer the data. No prompt or completion text is required.
What we need
Two pieces. Neither requires prompt or completion text or production write access.
easy / hard). CSV or JSONL, using the columns below.Trace columns
Use these header names. Column order does not matter. CSV and JSONL are both fine. Additional operational columns are okay, but do not include prompt or completion text.
| Field | Status | What to provide | Example |
|---|---|---|---|
timestamp_ms |
Required | Arrival time in milliseconds. Unix epoch ms or milliseconds from the start of the file both work. Rows do not need to be pre-sorted. | 0 or 1715313600123 |
difficulty_tier |
Required | Request eligibility for the standard first-pass audit: easy for requests eligible for the small-model pool and hard for requests that require the large-model pool. If you don’t already have eligibility labels, we’ll work with you to establish them. |
easy |
request_id |
Optional | A stable ID for the request. If omitted, we generate one from the row order. | req-000042 |
input_tokens |
Optional | Prompt/context token count. Not used by the standard first-pass audit. Include it if available; it is needed only for token-weighted analysis. | 256 |
output_tokens |
Optional | Generated token count. Not used by the standard first-pass audit. Include it if available; it is needed only for token-weighted analysis. | 64 |
Sample trace
The downloadable files match this layout. Empty optional cells are allowed.
| timestamp_ms | difficulty_tier | request_id | input_tokens | output_tokens |
|---|---|---|---|---|
| 0 | easy | req-0001 | ||
| 410 | hard | req-0002 | 512 | 180 |
| 980 | easy | req-0003 | ||
| 1500 | hard | req-0004 | 768 | 256 |
JSONL example
{"timestamp_ms": 0, "difficulty_tier": "easy", "request_id": "req-0001"}
{"timestamp_ms": 410, "difficulty_tier": "hard", "request_id": "req-0002", "input_tokens": 512, "output_tokens": 180}
Fleet note
The fleet note describes the inference resources that served the trace and how they are configured. The standard first-pass audit uses two serving pools: one small-model pool and one large-model pool. Each pool contains one model, with its own replica range, GPUs per replica, and serving capacity per replica.
If your system uses additional serving pools or multiple models within a pool, contact us. We can review your configuration with you and scope a targeted Efficiency Audit for your system.
| Item | Status | What to provide | Example |
|---|---|---|---|
| GPU platform and count | Required | GPU type and total number of GPUs available to the inference fleet. | 8 × NVIDIA A100 80GB |
| Serving pools | Required | List the model in each serving pool. | Small-model pool: Llama 3.2 3B Large-model pool: Llama 3.1 70B |
| Current capacity configuration | Required | For each pool, provide whether capacity is fixed or autoscaled, the replica count or autoscaling range, and the number of GPUs per replica. These limits are treated as deployment constraints in the audit. | Small-model pool: autoscaled, 1–4 replicas, 1 GPU/replica Large-model pool: autoscaled, 1–2 replicas, 2 GPUs/replica |
| Serving stack | Helpful | Serving system used by the fleet, such as vLLM, NVIDIA Dynamo, or SGLang. | vLLM |
| Service objectives | Helpful | Latency, throughput, error rate, or other service objectives, if available. | TTFT p99 < 1 s |
| Current routing behavior | Helpful | How requests are currently assigned between the two serving pools. | Router sends simpler requests to the small-model pool and more complex requests to the large-model pool. |
| Trace window | Helpful | Time period covered by the request trace. | Sept. 2, 2:00–3:00 PM ET |
| Per-replica serving capacity | Helpful | Typical request rate per replica for each pool, if known. If unknown, we’ll use a first-pass assumption and can refine it with you. | Small-model pool: ~6 requests/s per replica Large-model pool: ~3 requests/s per replica |
Example fleet note
Required GPU platform and count: 8 × NVIDIA A100 80GB Serving pools: Small-model pool: Llama 3.2 3B Large-model pool: Llama 3.1 70B Current capacity configuration: Small-model pool: autoscaled, 1–4 replicas, 1 GPU/replica Large-model pool: autoscaled, 1–2 replicas, 2 GPUs/replica Helpful Serving stack: vLLM Service objectives: TTFT p99 < 1 s; error rate < 1% Current routing behavior: Router sends simpler requests to the small-model pool and more complex requests to the large-model pool. Trace window: September 2, 2026, 2:00–3:00 PM ET Per-replica serving capacity (if known): Small-model pool: ~6 requests/s per replica Large-model pool: ~3 requests/s per replica
Download the fill-in template, or copy the example above and replace the values with yours.
Practical guidance
A representative 30–60 minute window of production-like traffic is a good starting point; longer traces are welcome. Use UTF-8 CSV or JSONL. There is no hard minimum request count.
The standard first-pass audit uses request timing and eligibility information — not prompt or completion content. Remove prompt, response, and other raw-text fields before sharing the trace.
Request an Efficiency Audit and we’ll review your use case and arrange a suitable way to share the required data. For first contact, please do not send production traces by email.
