Schedatrix Efficiency Audit - Fleet Note Share this with the sanitized request trace after a transfer method is agreed. Do not include prompt or completion text. The standard first-pass audit uses two serving pools: one small-model pool and one large-model pool. Each pool contains one model. Required GPU platform and count GPU type and total number of GPUs available to the inference fleet. Example: 8 x NVIDIA A100 80GB GPU platform and count: Serving pools List the model in each serving pool. Example: Small-model pool: Llama 3.2 3B Large-model pool: Llama 3.1 70B Serving pools: Small-model pool: Large-model pool: Current capacity configuration For each pool, provide whether capacity is fixed or autoscaled, the replica count or autoscaling range, and the number of GPUs per replica. These limits are treated as deployment constraints in the audit. Example: Small-model pool: autoscaled, 1-4 replicas, 1 GPU/replica Large-model pool: autoscaled, 1-2 replicas, 2 GPUs/replica Current capacity configuration: Small-model pool: Large-model pool: Helpful Serving stack Serving system used by the fleet, such as vLLM, NVIDIA Dynamo, or SGLang. Example: vLLM Serving stack: Service objectives Latency, throughput, error rate, or other service objectives, if available. Example: TTFT p99 < 1 s Service objectives: Current routing behavior How requests are currently assigned between the two serving pools. Example: Router sends simpler requests to the small-model pool and more complex requests to the large-model pool. Current routing behavior: Trace window Time period covered by the request trace. Example: Sept. 2, 2:00-3:00 PM ET Trace window: Per-replica serving capacity Typical request rate per replica for each pool, if known. If unknown, we'll use a first-pass assumption and can refine it with you. Example: Small-model pool: ~6 requests/s per replica Large-model pool: ~3 requests/s per replica Per-replica serving capacity: Small-model pool: Large-model pool: Have a more complex configuration? If your system uses additional serving pools or multiple models within a pool, contact us. We can review your configuration with you and scope a targeted Efficiency Audit for your system.