CATALOG

Inference endpoints. Pick what you need.

OpenAI-compatible inference for the latest open-source models. Hoonify routes your workload to capacity automatically.

Filters
Provider
Thinkingmachines

Inkling-Small

thinkingmachines/Inkling-Small
$0.50
/ 1M IN
$1.20
/ 1M OUT
$0.10 / 1M CACHED

Inkling-Small is an open-weight, general-purpose multimodal model that processes text, image, and audio inputs to generate text outputs. It is designed for developers to build AI applications, coding assistants, and chatbots, offering support for multiple natural and coding languages. 256K Context length.

Zai-org

GLM-5.2

zai-org/GLM-5.2
$1.40
/ 1M IN
$4.40
/ 1M OUT
$0.18 / 1M CACHED

GLM 5.2: Z AI's flagship model for long-horizon tasks. Mixture-of-Experts model with 744B parameters (40B active) and up to 1M-token context window.

Google

Gemma-4-31B-it

google/gemma-4-31B-it
$0.12
/ 1M IN
$0.38
/ 1M OUT
$0.09 / 1M CACHED

Gemma 4: 31 billion parameter dense multi-model model for deep reasoning and complex tasks. Context Length: 256K tokens. Language Support: 140 languages and dialects

Qwen

Qwen3.6-27B

Qwen/Qwen3.6-27B
$0.32
/ 1M IN
$3.20
/ 1M OUT
$0.15 / 1M CACHED

Qwen3.6: 27 Billion Parameter Dense Causal Language Model with Native Vision Encoder. Context Length: 262K tokens. Language Support: 201 languages and dialects