Skip to main content
L14.2

Inference APIs

Goal

Turn a text-generation service into an API a small client can call predictably: known fields, safe limits, useful errors, and changes that do not surprise existing clients.

Why did the homework app break?​

Suppose a homework app calls a text-generation service like this:

{"input":"Summarize photosynthesis in three sentences.","max_tokens":128}

Yesterday the server returned:

{"text":"Plants use light energy ...","finish_reason":"stop"}

Today someone improves the server and renames input to prompt and text to output. The model may be better, but the homework app now fails because the two programs no longer agree on the message shape.

A second client creates a different problem:

{"input":"Write a short answer.","max_tokens":4096}

The request is syntactically valid JSON, but suppose this service allows at most 512 output tokens. A request for 4096 is outside that contract and should be rejected before inference.

These are API problems rather than model problems. The client and server need a stable agreement. It says which fields exist, which types and ranges are allowed, what comes back, and how failures are reported. That agreement is the API contract. A schema is a machine-readable description of that shape.

Make one small request predictable​

Start with the smallest request that expresses the job. For a simple generation endpoint, the public fields might be:

{
"input": "Summarize photosynthesis in three sentences.",
"max_tokens": 128,
"temperature": 0.2
}

The service can define input as required text, max_tokens as an integer within a supported range, and temperature as an optional number with a documented default.

A useful response keeps the generated text separate from information the client may need for debugging:

{
"request_id": "r-17",
"text": "Plants use light energy ...",
"finish_reason": "stop",
"output_tokens": 48,
"latency_ms": 340,
"deployment": "school-helper-v3"
}

Now both sides know what to send and what to read. Fields such as output_tokens and latency_ms add small, stable measurements. The client can use them for debugging or monitoring without exposing the scheduler's internal objects.

Reject impossible work before inference​

Limits such as maximum input length and maximum output length protect the service as well as the model. Checking them before scheduling expensive work prevents one request from consuming an unreasonable share of memory or device time.

The same idea can apply to account quotas or supported model choices. Rejecting a request early is cheaper and easier to explain than allowing it to fail after the model begins running.

Return errors a client can act on​

Different failures need different responses. A client can fix malformed input, wait and retry after overload, ask for a supported model, or alert an operator when repeated internal failures occur.

A simplified set might look like:

400 invalid_request → fix the request
401 unauthorized → authenticate correctly
429 overloaded → back off and retry later
503 unavailable → service cannot currently run the request

The exact status-code design depends on the API, but the important lesson is stability: do not make every failure look like the same vague “generation failed” message.

OpenAPI is one common way to describe HTTP request and response shapes so tools can read the same rules humans document.

Change the interface without surprising old clients​

Adding an optional response field is usually less disruptive than renaming a required request field. If a change would alter the meaning of an existing field or remove something clients depend on, the service may need a new API version or a planned migration.

Request IDs also help when something changes. The client can report r-17, and operators can connect that one request to logs, traces, queue time, and the deployment version that handled it.

Keep internal knobs internal​

A serving runtime may support dozens of scheduler and decoding settings. Publishing every internal switch would make clients depend on implementation details that may change later.

Expose the controls clients genuinely need. Keep memory tuning, scheduler experiments, and most runtime-specific options inside the service. A narrow public API gives the implementation room to improve without breaking every caller.

Predict

A serving engine supports fifty tuning flags, but clients only need input text and a bounded output length. What should the public API expose?

Run the local Lab​

Run:

python3 labs/notebooks/level-14/l14-02-inference-api.py

The Lab validates request dictionaries before mock inference begins.

  1. Run it unchanged. The request uses max_tokens: 128, returns status 200, and records one mock inference call.
  2. The configured service limit is 512. Before editing, predict both the error category and inference-call count if only the request asks for 4096 output tokens.
  3. Change only "max_tokens": 128 in the main request to "max_tokens": 4096, then rerun.
  4. Confirm the response is a 400-class validation error with output_limit_exceeded and the main request does not reach mock inference.

Loading lab…

Quick Check

1. Why enforce output-token limits at the API boundary?
2. What is a useful inference response field besides generated text?
3. Why avoid exposing every runtime tuning flag publicly?

0 of 3 questions answered.

Explain it back​

Design a generate endpoint with request fields, bounds, response metadata, and four error categories. Explain which fields are client contract versus internal runtime configuration.

Key Takeaways

  • Inference APIs should have explicit schemas and bounds.
  • Stable error categories guide client behavior.
  • Include request and deployment identity in responses.
  • Machine-readable API descriptions reduce hidden assumptions.
  • Keep the public contract smaller than the serving runtime.

Next Lesson

Next, L14.3 — Batching and Streaming shows how the service shares compute while controlling first-response latency.

References

Lesson actions

Completion is stored locally on this device.

View progress