AI Model Hub

Every model on EcoLink — video, image, speech and language, self-hosted and partner-served alike. Pick one and start creating.

30 of 30 models
Price on request

Seedance 2.0

Text to Video
Image to Video
+2
From $0.05/sec

Wan2.2-T2V-A14B

Text to Video
From $0.05/sec

Wan2.2-S2V-14B (Presenter / Talking Head)

Speech-to-video: animates a portrait to speak driving audio.

Text to Video
Image to Video
+1
From $0.05/sec

LTX 2.5

Text-to-video with synchronized audio. Generates a soundtrack alongside the picture and returns a single MP4 with both tracks.

Text to Video
Image to Video
From $0.055/sec

Wan 3.0

Text to Video
Image to Video
+2
From $0.075/sec

wan3.0-video-prime

Text to Video
Image to Video
+2
Price on request

Seedance 2.5

Text to Video
Image to Video
+2
$0.01/pic

Z-Image-Turbo

Z-Image is a powerful and highly efficient image generation model family with 6B parameters. Currently there are four variants: 🚀 Z-Image-Turbo – A distilled version of Z-Image that matches or exceeds leading competitors with only 8 NFEs (Number of Function Evaluations). It offers ⚡️sub-second inference latency⚡️ on enterprise-grade H800 GPUs and fits comfortably within 16G VRAM consumer devices. It excels in photorealistic image generation, bilingual text rendering (English & Chinese), and robust instruction adherence. 🎨 Z-Image – The foundation model behind Z-Image-Turbo. Z-Image focuses on high-quality generation, rich aesthetics, strong diversity, and controllability, well-suited for creative generation, fine-tuning, and downstream development. It supports a wide range of artistic styles, effective negative prompting, and high diversity across identities, poses, compositions, and layouts. 🧱 Z-Image-Omni-Base – The versatile foundation model capable of both generation and editing tasks. By releasing this checkpoint, we aim to unlock the full potential for community-driven fine-tuning and custom development, providing the most "raw" and diverse starting point for the open-source community. ✍️ Z-Image-Edit – A variant fine-tuned on Z-Image specifically for image editing tasks. It supports creative image-to-image generation with impressive instruction-following capabilities, allowing for precise edits based on natural language prompts.

Text to Image
$0.02/pic

FLUX.2 Klein

The FLUX.2 [klein] model family are our fastest image models to date. FLUX.2 [klein] unifies generation and editing in a single compact architecture, delivering state-of-the-art quality with end-to-end inference in as low as under a second. Built for applications that require real-time image generation without sacrificing quality. FLUX.2 [klein] 9B is a 9 billion parameter rectified flow transformer capable of generating images from text descriptions and supports multi-reference editing capabilities. Our flagship small model. Defines the Pareto frontier for quality vs. latency across text-to-image, single-reference editing, and multi-reference generation. Matches or exceeds models 5x its size—in under half a second. Built on a 9B flow model with 8B Qwen3 text embedder, step-distilled to 4 inference steps.

Text to Image
Image to Image
$0.03/pic

Qwen-Image

Qwen-Image, an image generation foundation model in the Qwen series that achieves significant advances in complex text rendering and precise image editing. Experiments show strong general capabilities in both image generation and editing, with exceptional performance in text rendering, especially for Chinese.

Text to Image
$0.006/min

Kokoro-82M

Fast, lightweight text-to-speech model

Text to Speech
$0.003/min

Qwen3-ASR-1.7B

The Qwen3-ASR family includes Qwen3-ASR-1.7B and Qwen3-ASR-0.6B, which support language identification and ASR for 52 languages and dialects. Both leverage large-scale speech training data and the strong audio understanding capability of their foundation model, Qwen3-Omni. Experiments show that the 1.7B version achieves state-of-the-art performance among open-source ASR models and is competitive with the strongest proprietary commercial APIs.

Speech to Text
$0.012/min

Qwen3-TTS

Qwen3-TTS covers 10 major languages (Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, and Italian) as well as multiple dialectal voice profiles to meet global application needs. In addition, the models feature strong contextual understanding, enabling adaptive control of tone, speaking rate, and emotional expression based on instructions and text semantics, and they show markedly improved robustness to noisy input text.

Text to Speech
$0.003/min

Fun-ASR-Nano

LLM-Powered Speech Recognition — 31 Languages, Dialects & Accents End-to-end ASR trained on tens of millions of hours of data. Supports Chinese (+ dialects), English, Japanese, Korean, French, German, Spanish, and 24 more languages.

Speech to Text
$0.003/min

ViiTorVoice-NAR

ViiTorVoice-NAR is a non-autoregressive speech generation model for voice cloning, local speech editing, and emotion / paralinguistic speech control.

Voice Cloning
$0.006/min

Whisper-Large-V3-Turbo

Whisper large-v3-turbo is a finetuned version of a pruned Whisper large-v3. In other words, it's the exact same model, except that the number of decoding layers have reduced from 32 to 4. As a result, the model is way faster, at the expense of a minor quality degradation.

Speech to Text
$0.50/1M tok

Gemma-4-31B-IT

Gemma 4 models are designed to deliver frontier-level performance at each size, targeting deployment scenarios from mobile and edge devices (E2B, E4B) to consumer GPUs and workstations (26B A4B, 31B). They are well-suited for reasoning, agentic workflows, coding, and multimodal understanding. The models employ a hybrid attention mechanism that interleaves local sliding window attention with full global attention, ensuring the final layer is always global. This hybrid design delivers the processing speed and low memory footprint of a lightweight model without sacrificing the deep awareness required for complex, long-context tasks. To optimize memory for long contexts, global layers feature unified Keys and Values, and apply Proportional RoPE (p-RoPE).

Vision
$0.80/1M tok

qwen3-omni-30b-a3b-instruct

Vision
$0.50/1M tok

qwen3-vl-8b-instruct

Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date. This generation delivers comprehensive upgrades across the board: superior text understanding & generation, deeper visual perception & reasoning, extended context length, enhanced spatial and video dynamics comprehension, and stronger agent interaction capabilities. Available in Dense and MoE architectures that scale from edge to cloud, with Instruct and reasoning‑enhanced Thinking editions for flexible, on‑demand deployment. Key Enhancements: Visual Agent: Operates PC/mobile GUIs—recognizes elements, understands functions, invokes tools, completes tasks. Visual Coding Boost: Generates Draw.io/HTML/CSS/JS from images/videos. Advanced Spatial Perception: Judges object positions, viewpoints, and occlusions; provides stronger 2D grounding and enables 3D grounding for spatial reasoning and embodied AI. Long Context & Video Understanding: Native 256K context, expandable to 1M; handles books and hours-long video with full recall and second-level indexing. Enhanced Multimodal Reasoning: Excels in STEM/Math—causal analysis and logical, evidence-based answers. Upgraded Visual Recognition: Broader, higher-quality pretraining is able to “recognize everything”—celebrities, anime, products, landmarks, flora/fauna, etc. Expanded OCR: Supports 32 languages (up from 19); robust in low light, blur, and tilt; better with rare/ancient characters and jargon; improved long-document structure parsing. Text Understanding on par with pure LLMs: Seamless text–vision fusion for lossless, unified comprehension.

Vision
$0.60/1M tok

Qwen3.6-27B

Qwen3.6 27B dense vision-language model. Text + image input, 32K context, FP8. Supports tool/function calling and reasoning.

Vision
$0.40/1M tok

Qwen3.6-35B-A3B

Qwen3.6 35B-A3B MoE vision-language model (35B total / 3B active). Text + image input, 32K context, FP8. Supports tool/function calling and reasoning.

Vision
$0.60/1M tok

Qwen3.8-27B

Qwen3.8 27B dense vision-language model. Text + image input, 32K context, FP8. Supports tool/function calling and reasoning.

Vision
$1.42/1M tok

DeepSeek-V4-Flash

DeepSeek-V4-Flash with 284B parameters (13B activated) — both supporting a context length of one million tokens.

Language
$2.72/1M tok

DeepSeek-V4-Pro

Language
$3.00/1M tok

GLM-5.2

GLM-5.2 marks a substantial leap in long-horizon task capability over its predecessor GLM-5.1 and, for the first time, delivers that capability on a solid 1M-token context. GLM-5.2's new capabilities include: Solid 1M Context: A solid 1M-token context that stably sustains long-horizon work Advanced Coding with Flexible Effort: Stronger coding capabilities with multiple thinking effort levels to balance performance and latency Improved Architecture: We propose IndexShare, which reuses the same indexer across every four sparse attention layers, reducing per-token FLOPs by 2.9× at a 1M context length. We also improve GLM-5.2’s MTP layer for speculative decoding, increasing the acceptance length by up to 20% Pure Open: An MIT open-source license — no regional limits, technical access without borders

Language
$0.30/1M tok

qwen3-coder-30b-a3b-instruct

Qwen3-Coder-30B-A3B-Instruct has the following features: Type: Causal Language Models Training Stage: Pretraining & Post-training Number of Parameters: 30.5B in total and 3.3B activated Number of Layers: 48 Number of Attention Heads (GQA): 32 for Q and 4 for KV Number of Experts: 128 Number of Activated Experts: 8 Context Length: 262,144 natively.

Language
$12.00/1M tok

Kimi-K3

Language
$4.30/1M tok

GLM-5.3

Language
$0.90/1M tok

MiniMax-M3

Language
$5.70/1M tok

qwen3.8-max

Language