Provider API specification¶
API overview¶
Examples containing await run inside an async function. Choose models available to your account; capabilities depend on the provider, model, and API format.
Basic usage¶
import republic
from republic.events import TextDelta
model = republic.get_model("openai:gpt-6-sol")
# Chat
response = await model.chat("Hello, how are you?")
print(response.text)
# Stream
async with model.stream("Hello, how are you?") as stream:
async for event in stream:
if isinstance(event, TextDelta):
print(event.chunk)
print(stream.response.text)
Consume the stream before reading its final response or shortcuts such as stream.output. Reading them early raises errors.StreamNotFinishedError.
A clean EOF alone does not complete a stream. Chat Completions requires [DONE]; Responses requires response.completed, response.incomplete, or the [DONE] marker used by compatible services; Messages requires message_stop (Copilot also accepts [DONE]); Gemini requires a candidate's finishReason. A provider-reported output limit or paused turn is a terminal result, not a transport truncation.
If the stream ends without its completion signal, Republic raises errors.StreamIncompleteError, emits no Completed event and writes no conversation history. Any partial events already yielded remain available to the caller, but stream.response stays unavailable. Transport failures and provider error events also leave history unchanged. Custom StreamParser implementations must set self.completed = True on their terminal event, or override finish() to validate completion according to their protocol.
Initialization¶
# Use the API key and base URL from REPUBLIC_OPENAI_* environment variables.
model = republic.get_model("openai:gpt-6-sol")
# Change the environment variable prefix.
provider = republic.get_provider("openai", env_prefix="MY_CUSTOM_PREFIX")
model = provider.get_model("gpt-6-sol")
# Pass an API key and base URL loaded by the application.
model = republic.get_model("openai:gpt-6-sol", api_base=api_base, api_key=api_key)
# Pass an application-defined HTTPX2-compatible auth object.
model = republic.get_model("openai:gpt-6-sol", auth=MyAuth())
# Use a provider instance.
provider = republic.get_provider("openai", api_key=api_key)
model = provider.get_model("gpt-6-sol")
Here, api_base, api_key, and MyAuth are supplied by the application. A top-level constructor accepts provider:model; a provider instance accepts only the model name. See configuration for options and credential precedence, and authentication for API keys and account logins.
HTTP client lifecycle¶
A Provider lazily creates one HTTP client and reuses it for requests, streams, model listing, and every model created by that Provider. Separate calls to the top-level model factories create separate Providers.
For applications that manage a long-lived Provider, call await provider.close() at shutdown:
provider = republic.get_provider("openai")
model = provider.get_model("gpt-6-sol")
try:
response = await model.chat("Hello")
finally:
await provider.close()
Providers also support async context management. Context entries are released on exit, including when the body raises an exception or is cancelled:
async with republic.get_provider("openai") as provider:
first = provider.get_model("gpt-6-sol")
second = provider.get_model("gpt-6-luna")
await first.chat("Hello")
await second.chat("Hello")
Chat, embedding, and decision models support the same pattern. The Provider tracks entered Provider and model contexts and closes only when all of them have exited:
Use the same pattern with get_embedding_model() and get_decision_model(), or enter an existing model with async with model:. Entering a Provider or model does not create the HTTP client before its first request. Creating a model without entering its context does not register it as open. Nested contexts for the same model remain open until all its entries have exited.
provider = republic.get_provider("openai")
async with provider.get_model("gpt-6-sol") as first:
async with provider.get_embedding_model("text-embedding-3-small") as embeddings:
vector = await embeddings.embed("Hello")
# The Provider remains open while first's context is active.
response = await first.chat("Hello")
# The last model context has exited, so the Provider is now closed.
A Provider context keeps the client open while model contexts enter and exit, so sequential model contexts can share it:
async with republic.get_provider("openai") as provider:
for name in ["gpt-6-sol", "gpt-6-luna"]:
async with provider.get_model(name) as model:
response = await model.chat("Hello")
# The Provider context still keeps the client open here.
Nested Provider contexts follow the same rule: exiting an inner context does not close the client while another Provider or model context is active. Explicit await provider.close() closes it immediately, regardless of open contexts.
provider.close() is idempotent; provider.is_closed reports whether it has been closed. Once closed, all models sharing that Provider reject further requests, and the Provider cannot be reopened. Finish outstanding requests and exit stream contexts before closing it. Exiting a stream context closes only that response and keeps the HTTP client available for subsequent requests.
An internally created client must be used and closed on the event loop that first used it. Constructing the Provider outside an event loop is supported; reusing an active Provider across multiple asyncio.run() calls raises RuntimeError.
An external client passed through http_client= remains caller-owned: closing the Provider leaves that client open. See HTTP client configuration for transport settings.
Tool calls¶
tool1 and tool2 below are republic.Tool schemas.
Each call has an ID, name, raw JSON arguments, decoded args, and provider metadata. Preserve the call when returning its result.
Structured input¶
response = await model.chat([
republic.system("You are a helpful assistant."),
republic.user("Hello, how are you?"),
])
print(response.text)
Multimodal content¶
User messages accept text and Image, Audio, and Video parts. The public content shape is independent of the selected provider:
| Part | Constructor | Source fields |
|---|---|---|
Image |
republic.image(source, media_type=...) |
media_type, data or url |
Audio |
republic.audio(source, media_type=...) |
media_type, data or url |
Video |
republic.video(source, media_type=...) |
media_type, data or url |
media_type is a MIME type such as audio/mpeg, rather than a protocol's format label such as mp3. data holds raw bytes; url holds a remote reference. The loader helpers handle sources consistently:
- A local path or
Pathloads bytes and infers the MIME type from the filename unlessmedia_type=is supplied. - An HTTP(S) or
gs://URL keeps a reference without downloading it. MIME inference uses the URL path; query strings do not change its type. Supplymedia_type=for URLs without a recognizable filename. - A base64 data URL decodes bytes and uses the MIME type declared in that URL.
- Raw bytes require an explicit
media_type=.
Use the parts directly in a user message. Each API format owns its wire encoding:
model = republic.get_model("google:MODEL_ID")
response = await model.chat(
republic.user(
"Describe these inputs.",
republic.image("path/to/image.png"),
republic.audio("path/to/audio.wav"),
republic.video("path/to/video.mp4"),
)
)
print(response.text)
The same message works with model.stream(). The following table describes Republic's input encoders; the chosen service and model must also support that input.
| API format | Images | Audio | Video |
|---|---|---|---|
| Chat Completions | Data or URL | Inline WAV or MP3 | Data or URL, when the service supports video_url |
| Responses | Data or URL | Rejected | Rejected |
| Messages | Data or URL | Rejected | Rejected |
| Gemini | Inline data or file reference | Inline data or file reference | Inline data or file reference |
Provider dialects may extend a format. For example, OpenRouter's Chat Completions encoder also maps OGG, FLAC, and its other supported MIME types to their protocol format labels. Custom Chat Completions dialects can override audio_content() or extend AUDIO_FORMATS. Application code still passes the same Audio part.
Unsupported audio combinations raise errors.UnsupportedFeatureError before sending a request. Republic does not transcode audio or fetch remote references to make them inline.
After executing tool calls from a response, return the results with the full assistant message. Here, results contains republic.tool(call, output) messages created by the application:
answer = await model.chat([
"Hello, how are you?",
response.message,
*results,
])
print(answer.text)
response.message preserves tool IDs and opaque provider state. See tool round trips for a complete example.
Structured output¶
Pass a Pydantic model or another type Pydantic can validate as output_schema. MyOutputSchema below is defined by the application:
response = await model.chat("Hello, how are you?", output_schema=MyOutputSchema)
print(response.output)
async with model.stream("Hello, how are you?", output_schema=MyOutputSchema) as stream:
async for event in stream:
if isinstance(event, TextDelta):
print(event.chunk)
print(stream.output)
The completed result is validated from the response's JSON text. A refusal leaves output as None; invalid output raises a Pydantic validation error. See structured output.
Image output¶
Request image generation using a provider-run tool and a model that supports it:
model = republic.get_model("openai:gpt-6-sol")
response = await model.chat("Generate an image of a tree.", tools=[republic.tools.ImageGeneration()])
print(response.text)
print(response.image_parts)
response.image_parts contains republic.Image values. With republic[image] installed, response.images decodes inline image data into Pillow images. The response API has no video-output accessor.
Token usage¶
response = await model.chat("Hello, how are you?")
print(response.token_usage)
print(response.token_usage.total_tokens)
Total tokens are input plus output tokens. Reasoning tokens are included in output; cached and cache-write tokens are included in input when reported.
Providers¶
Built-in providers¶
The registered names are openai, anthropic, google, openrouter, typesafe, codex, github-copilot, grok, azure-openai, ollama, deepseek, moonshot, zai, together, mistral, minimax, and magpie. republic.all_providers() returns every registered name, including custom providers, sorted. See supported providers for model kinds, formats, and authentication paths.
Listing models¶
provider = republic.get_provider("openai")
for model in await provider.list_models():
print(model.id, model.display_name)
list_models() returns republic.ModelInfo values for every page of the service's model list. Pass model.id to get_model() or another model getter; model.raw keeps the service's entry, such as context length or pricing. Codex has no model list and raises errors.UnsupportedFeatureError.
A custom provider sets MODELS_PATH relative to api_base, or None when the service has no list. Override _parse_models() for another response shape and _models_params() for pagination.
Custom providers¶
Subclass a provider and register a unique name. Set the endpoint and formats to match the service:
from republic.providers import OpenAI
class MyCustomProvider(OpenAI):
DEFAULT_API_BASE = "https://api.example.com/v1"
SUPPORTED_API_FORMATS = ("responses", "messages", "chat")
republic.register_provider(MyCustomProvider, "custom")
model = republic.get_model("custom:gpt-6-sol")
Service dialects¶
Services that speak a known format with their own fields plug in through format hooks. Subclass a format from republic.formats and set it as CHAT_FORMAT on an OpenAICompatible subclass, or return it from Provider.select_api_format(), which also receives the model name:
from republic.formats import ChatFormat
from republic.providers import OpenAICompatible
class AcmeChat(ChatFormat):
def reasoning_fields(self, effort, *, include_reasoning):
return {"thinking": {"type": "disabled" if effort == "none" else "enabled"}} if effort else {}
class Acme(OpenAICompatible):
name = "acme"
DEFAULT_API_BASE = "https://api.acme.example/v1"
CHAT_FORMAT = AcmeChat()
republic.register_provider(Acme)
Every chat format has reasoning_fields(). ChatFormat also has max_tokens_fields(), reasoning_text(), content_text() for answer text in content parts, and assistant_fields() for fields added to assistant messages sent back, such as reasoning_content.
API format support¶
Republic selects the first format of the requested model kind in the provider's SUPPORTED_API_FORMATS. An explicit api_format= takes precedence for formats of that kind. Each built-in provider's default is the first chat format in the provider directory.
Use api_format= to select a supported format explicitly:
model = republic.get_model("openai:gpt-6-sol", api_format="chat")
provider = republic.get_provider("openai", api_format="chat")
model = provider.get_model("gpt-6-sol")
A format not supported by the provider, such as messages for OpenAI, raises errors.UnsupportedApiFormatError. A preferred chat format does not prevent selecting an embedding or decision format for another model kind.
Optional conversation history¶
from republic.history import InMemoryHistory
model = republic.get_model("openai:gpt-6-sol", history=InMemoryHistory(max_entries=10))
response = await model.chat("Hello, how are you?")
print(response.text)
response = await model.chat("How about you?")
print(response.text)
History follows this protocol, exported from republic.history:
from typing import Protocol
from republic import Message
class HistoryProtocol(Protocol):
async def read(self) -> list[Message]:
"""Read the conversation history."""
...
async def write(self, messages: list[Message]) -> None:
"""Append messages to the conversation history."""
...
With attached history, Republic reads earlier messages before each request and appends the new input and assistant response after completion. Without it, the caller supplies conversation history.
Non-chat models¶
Embedding models return vectors:
model = republic.get_embedding_model("openai:embedding-model")
response = await model.embed("Hello, how are you?")
print(response.vector)
embed_many(texts) returns one vector per input in response.vectors. Replace embedding-model with an available embedding model.
Decision models answer typed questions about state:
from republic.decisions import Noul
model = republic.get_decision_model("typesafe:decision-model")
response = await model.decide(
"The weather is clear and I have an hour free.",
questions={"decision": Noul("Should I go for a walk?")},
)
print(response.answers["decision"])
Replace decision-model with an available decision model. Choice, Noul, and Score describe the questions; answers use the same IDs. Each provider exposes only the model kinds it supports.
Source: Frost Ming's Republic specification.