Specialized AI models. One API.
Access fast and cost-efficient ASR, TTS, OCR and vision models through a unified API or MCP.
One credential · Typed tasks · Structured outputs
Capabilities
Purpose-built for the work between prompts.
Give each task the model, interface and output shape it actually needs.
Speech to Text
Text to Speech
OCR
Object Detection
Unified access
One API, many tasks
Add speech, document and vision capabilities without rebuilding authentication, usage controls or operational plumbing for every model.
Elastic inference lets task containers start and scale with demand, reducing idle infrastructure when traffic is low.
Two ways in
Built for agents
Applications and agents reach the same task layer. Choose the interface that fits the caller, not a different product.
For applications and automation
Direct requests · predictable responses
For agents and tool platforms
Tool discovery · structured invocation
A clear task contract, whatever the model.
Send a typed task and input asset. Receive a structured result that your application or agent can use without translating free-form output.
This request shape illustrates the interface direction. Verified endpoint examples will ship with the public quickstart.
TASK REQUEST
authorization Bearer $TASKINFER_API_KEY
task speech.transcribe
input <audio asset>
STRUCTURED RESULT
status completed
output <task-specific result>Right-sized inference
Why specialized models
General-purpose models remain useful for open-ended reasoning. For repeatable production tasks, a focused model can offer a simpler, more controllable operating profile.
Start with a task
Use the right model for the task.
Add focused speech, document and vision capabilities through one inference layer built for developers and agents.
