# Inputs and modalities

> How an LLM Surface accepts three client API formats, opts in to audio, image, and realtime voice input, checks and extracts attached documents, and routes requests by input type.

Clients send requests in different shapes. One application uses the OpenAI SDK and another the Anthropic SDK, while others attach images, audio files, or documents that only some models can read. An LLM Surface accepts all three supported client API formats by default, and accepts audio, images, and realtime voice only when you add the matching Inputs element to its canvas. A Document File element checks attached documents and can turn them into prompt text that any model can read. [Server tools and modalities →](/products/affinidi-trust-fabric/agent-stream/reference/surfaces/server-tools-and-modalities.md)

## How a surface answers three client API formats

A client chooses its request format through the endpoint path it calls on the surface’s route:

| Path | Client API format |
| /v1/chat/completions | OpenAI Chat Completions |
| /v1/messages | Anthropic Messages |
| /v1/responses | OpenAI Responses |

The surface converts each request into one internal shape before its pipeline runs, and returns the response in the format the client used. The format a client speaks is independent of the surface’s provider, so an Anthropic SDK can call a surface backed by an OpenAI model.

Every LLM Surface accepts all three formats by default. You can turn one off in the Client API formats section of the surface’s configuration editor, opened with Configure Surface Features. A request in a format the surface has turned off gets an HTTP 404 response with an error naming the format the surface doesn’t accept.

## How Inputs elements opt a surface in to attachments

The Inputs palette group holds one element per kind of input. Dropping an element onto an LLM Surface canvas switches that input on, or for Document File, switches on document checks and extraction. Each element’s sidebar has an enable switch. A disabled element behaves the same as no element.

| Element | What it accepts | Without the element |
| Audio File | Audio files attached to a chat message, in WAV or MP3, for a model that accepts audio input. | The request is rejected. |
| Image File | Images attached to a chat message, embedded in PNG, JPEG, GIF, or WebP format or given as an image URL. | The request is rejected. |
| Document File | Documents attached to a chat message, which the surface can turn into prompt text. | The document is passed to the provider as sent, without extraction or checks. |
| Audio Stream | Realtime voice sessions over a WebSocket connection. | The session is refused. |

Audio and image input are strictly opt-in. A request carrying audio or an image to a surface without an enabled matching element is rejected with an HTTP 400 error before it reaches the provider. With Audio File in place, audio is still rejected when the surface’s model doesn’t accept audio input, or when a file’s format or the total audio size falls outside the element’s settings. Image File checks the format and size of images embedded in the request and forwards an image URL as it is, because the appliance can’t inspect a remote image without fetching it.

The attachment types a request carries are also passed to the surface’s [Open Policy Agent (OPA) policy](/products/affinidi-trust-fabric/agent-stream/concepts/opa-policies.md), so a policy can allow an attachment type for some callers and deny it for others.

## How attached documents become prompt text

Through document ingestion, a surface with a Document File element reads a document attached to a chat message and turns it into ordinary prompt text. It first checks a document embedded in the request against the element’s allowed types and file size limit, and rejects the request with an HTTP 400 error if either check fails. A document given by reference, such as a provider file ID, is forwarded as it is. The element’s Mode then decides what happens to an accepted document:

- Auto, the default, extracts text from Text, Markdown, CSV, JSON, and DOCX files and forwards any other allowed type to the model as a file, such as a PDF when PDF is an allowed type.

- Extract extracts text from Text, Markdown, CSV, JSON, and DOCX files, and rejects the request for any allowed document it can’t extract, such as a PDF.

- Passthrough forwards every document as a file, for a model that reads documents natively. The allowed types and size limit still apply.

Extraction runs before [personally identifiable information (PII) protection](/products/affinidi-trust-fabric/agent-stream/concepts/pii-protection.md) and the request-side guardrails, so the extracted text is checked in the same way as text the caller typed. Text longer than the element’s character limit is truncated.

## How the Router sends a request to the right variant

The Router element, also in the Inputs group, holds an ordered list of routing rules. Each rule matches a Header value, a Body field value, or an Attachment type, and sends the request to the variant it names. The attachment types a rule can match are Image file, Audio file, Document file, and Realtime voice (WebSocket). The first matching rule wins. A request that matches no rule falls through to the canary split, and then to the surface’s default: the base surface, or the variant promoted to default. Rules apply while the Router’s Enforce routing switch is on, and are kept when you turn it off.

For a chat request, an explicit $alias in the route takes precedence over a routing rule, and a routing rule takes precedence over a canary split. See [Variants and progressive rollout](/products/affinidi-trust-fabric/agent-stream/concepts/variants-and-rollout.md) for the full selection order.

The variant is chosen before authentication, policy, and guardrails run. Each variant keeps its own copy of the settings it covers, including its model, its Inputs elements, and its Judge and Jury, so the chosen variant’s settings govern the rest of the request. This lets one route serve several kinds of input:

- Chat: plain text stays on the surface’s default and its text model.

- Images: an Image file rule sends requests with images to a variant with a vision-capable model and an Image File element.

- Voice: a Realtime voice (WebSocket) rule sends realtime sessions to a variant with a realtime model and an Audio Stream element.

A surface or variant whose model is a realtime model accepts only realtime sessions and rejects chat requests, which is why voice belongs on its own variant. For a realtime session, the WebSocket connection itself is the routing signal: while Enforce routing is on, the session goes to the variant named by the first Realtime voice (WebSocket) rule that names a variant. Otherwise it goes to the surface’s default.

## How realtime voice sessions are governed

With an enabled Audio Stream element, a caller opens a WebSocket connection to the surface’s /v1/realtime path, and the appliance relays live audio in both directions between the caller and the model. The surface’s model must be realtime-capable: an OpenAI realtime model or a Gemini Live model. The appliance refuses the session when the model isn’t realtime-capable.

Before a session opens, the caller is authenticated with the surface’s source authentication, and the session is checked against the appliance-wide policies and the surface’s own policy, as a chat request would be. Each session ends when it reaches its maximum session length.

Voice safety works on each turn of the conversation. With Review realtime model replies switched on in the Realtime response safety section, the surface’s [Jury](/products/affinidi-trust-fabric/agent-stream/concepts/guardrails.md) reviews the transcript of each model reply. This review needs the surface’s Jury to be enabled. The Enforcement setting decides how a reply reaches the caller:

- Gated holds each reply until the Jury approves it, and releases it only if it passes.

- Barge-in streams the reply live and cuts it off when the Jury flags it, so a fragment can play before the cut-off.

A blocked turn, or a Jury that fails to return a verdict, ends that reply with a refusal. If a refusal clip is selected in the Audio streams section, the clip plays. Otherwise the caller gets the Refusal message you set, or a default one.

The Judge doesn’t run on realtime voice, because caller audio reaches the model before any transcript exists for the Judge to review. The dashboard won’t save a canvas with both an enabled Judge and an enabled Audio Stream element, so voice safety relies on the Jury.

## Use cases

- Mixed SDK clients: teams using the OpenAI SDK and the Anthropic SDK call the same surface, each in its own format, under one set of guardrails and budgets.

- Document questions: a support assistant attaches DOCX policy documents to a question, and the surface extracts the text so that Prompt Guard can redact or pseudonymise personal data before the model sees it.

- Chat, vision, and voice on one route: the Router sends plain text to a text model, image requests to a vision-capable variant, and realtime sessions to a voice variant reviewed by the Jury.

## Related

- [Server tools and modalities](/products/affinidi-trust-fabric/agent-stream/reference/surfaces/server-tools-and-modalities.md): Field reference for the Inputs elements and realtime voice settings.

- [Routing and variants](/products/affinidi-trust-fabric/agent-stream/reference/surfaces/routing-and-variants.md): Field reference for routing rules, canary splits, and variants.

- [Surfaces](/products/affinidi-trust-fabric/agent-stream/concepts/surfaces.md): What an LLM Surface owns and how its canvas is composed.

- [Variants and progressive rollout](/products/affinidi-trust-fabric/agent-stream/concepts/variants-and-rollout.md): How variants are selected and promoted.

- [Guardrails](/products/affinidi-trust-fabric/agent-stream/concepts/guardrails.md): How the Judge and Jury review requests and responses.
