Schedatrix Efficiency Audit

Sample trace format

Send a sanitized request trace in the format below, plus a short fleet note. We’ll arrange a suitable way to transfer the data. No prompt or completion text is required.

What we need

Two pieces. Neither requires prompt or completion text or production write access.

1. Request traceWhen each request arrived, plus eligibility (easy / hard). CSV or JSONL, using the columns below.
2. Fleet noteWhat GPUs you have, two serving pools (one small-model pool and one large-model pool), and how those pools are provisioned today.

Trace columns

Use these header names. Column order does not matter. CSV and JSONL are both fine. Additional operational columns are okay, but do not include prompt or completion text.

FieldStatusWhat to provideExample
timestamp_ms Required Arrival time in milliseconds. Unix epoch ms or milliseconds from the start of the file both work. Rows do not need to be pre-sorted. 0 or 1715313600123
difficulty_tier Required Request eligibility for the standard first-pass audit: easy for requests eligible for the small-model pool and hard for requests that require the large-model pool. If you don’t already have eligibility labels, we’ll work with you to establish them. easy
request_id Optional A stable ID for the request. If omitted, we generate one from the row order. req-000042
input_tokens Optional Prompt/context token count. Not used by the standard first-pass audit. Include it if available; it is needed only for token-weighted analysis. 256
output_tokens Optional Generated token count. Not used by the standard first-pass audit. Include it if available; it is needed only for token-weighted analysis. 64

Sample trace

The downloadable files match this layout. Empty optional cells are allowed.

timestamp_ms difficulty_tier request_id input_tokens output_tokens
0easyreq-0001
410hardreq-0002512180
980easyreq-0003
1500hardreq-0004768256

JSONL example

{"timestamp_ms": 0, "difficulty_tier": "easy", "request_id": "req-0001"}
{"timestamp_ms": 410, "difficulty_tier": "hard", "request_id": "req-0002", "input_tokens": 512, "output_tokens": 180}

Fleet note

The fleet note describes the inference resources that served the trace and how they are configured. The standard first-pass audit uses two serving pools: one small-model pool and one large-model pool. Each pool contains one model, with its own replica range, GPUs per replica, and serving capacity per replica.

Have a more complex configuration?

If your system uses additional serving pools or multiple models within a pool, contact us. We can review your configuration with you and scope a targeted Efficiency Audit for your system.

ItemStatusWhat to provideExample
GPU platform and count Required GPU type and total number of GPUs available to the inference fleet. 8 × NVIDIA A100 80GB
Serving pools Required List the model in each serving pool. Small-model pool: Llama 3.2 3B
Large-model pool: Llama 3.1 70B
Current capacity configuration Required For each pool, provide whether capacity is fixed or autoscaled, the replica count or autoscaling range, and the number of GPUs per replica. These limits are treated as deployment constraints in the audit. Small-model pool: autoscaled, 1–4 replicas, 1 GPU/replica
Large-model pool: autoscaled, 1–2 replicas, 2 GPUs/replica
Serving stack Helpful Serving system used by the fleet, such as vLLM, NVIDIA Dynamo, or SGLang. vLLM
Service objectives Helpful Latency, throughput, error rate, or other service objectives, if available. TTFT p99 < 1 s
Current routing behavior Helpful How requests are currently assigned between the two serving pools. Router sends simpler requests to the small-model pool and more complex requests to the large-model pool.
Trace window Helpful Time period covered by the request trace. Sept. 2, 2:00–3:00 PM ET
Per-replica serving capacity Helpful Typical request rate per replica for each pool, if known. If unknown, we’ll use a first-pass assumption and can refine it with you. Small-model pool: ~6 requests/s per replica
Large-model pool: ~3 requests/s per replica

Example fleet note

Required
GPU platform and count: 8 × NVIDIA A100 80GB
Serving pools:
  Small-model pool: Llama 3.2 3B
  Large-model pool: Llama 3.1 70B
Current capacity configuration:
  Small-model pool: autoscaled, 1–4 replicas, 1 GPU/replica
  Large-model pool: autoscaled, 1–2 replicas, 2 GPUs/replica

Helpful
Serving stack: vLLM
Service objectives: TTFT p99 < 1 s; error rate < 1%
Current routing behavior: Router sends simpler requests to the small-model pool and more complex requests to the large-model pool.
Trace window: September 2, 2026, 2:00–3:00 PM ET
Per-replica serving capacity (if known):
  Small-model pool: ~6 requests/s per replica
  Large-model pool: ~3 requests/s per replica

Download the fill-in template, or copy the example above and replace the values with yours.

Practical guidance

What works well

A representative 30–60 minute window of production-like traffic is a good starting point; longer traces are welcome. Use UTF-8 CSV or JSONL. There is no hard minimum request count.

Do not include prompt or completion text

The standard first-pass audit uses request timing and eligibility information — not prompt or completion content. Remove prompt, response, and other raw-text fields before sharing the trace.

Ready to start?

Request an Efficiency Audit and we’ll review your use case and arrange a suitable way to share the required data. For first contact, please do not send production traces by email.