Fireworks AI Models
308 entities
MiMo-V2.6-Pro-RL
modelMiMo-V2.6-Pro-RL is a 1T-parameter sparse MoE omnimodal model (text, image, video, audio) with 1M-token context, designed as Xiaomi’s flagship agentic model. It scales a unified reinforcement-learning pipeline across coding, general agents, visual tasks, and cybersecurity using groupwise agentic grading and speculative decoding to drive self-improvement and strong benchmark performance
Ember-1
modelEmber-1 is a specialized model from Fireworks. Built on Kimi K3, it produces shorter reasoning traces, using approximately 40% fewer tokens while maintaining comparable quality across our evaluations
DeepSeek V4.1 Flash
modelDeepSeek-V4.1-Flash is a multimodal Mixture-of-Experts model with 552B backbone parameters that natively processes images and text at up to 1M token context. Its Causal Encoder-Decoder architecture activates only 8B parameters during prefill and 16B during decode, cutting the KV cache footprint to roughly a quarter of DeepSeek-V4-Flash for cost-efficient agentic workloads.
Ling 3 Flash Fin
modelLing-3.0-flash-Fin is the first finance-enhanced model in the Ant Ling family. Developed by Ant Group with leading financial institutions and domain experts, it extends Ling-3.0-flash through continued training on high-quality financial data. With 124B total parameters, 5.1B activated parameters, and a 256K context window, the model combines financial expertise with efficient inference for long-horizon agent workflows.
DeepSeek-V4-Flash-Vision-Exp
modelDeepSeek-V4-Flash-Vision-Exp is the first experimental multimodal model in the DeepSeek-V4 family. It builds on DeepSeek-V4-Flash by adding visual modules and continued training for visual understanding, with substantially improved multimodal agent capabilities while remaining comparable on text-only agent tasks. It has the same model structure as DeepSeek-V4-Flash-DSpark, i.e. it comes with a speculative decoding module attached.
GLM 5.3 Flash
modelGLM-5.3-Flash is the first natively multimodal model in the GLM-5 series with 320B total parameters and 18B active parameters. It incorporates several architectural improvements over GLM-5.2 including a hybrid architecture, sharply reducing long-context serving costs while preserving precise long-context and Manifold-Constrained Hyper-Connections (mHC) to further improve scaling efficiency.
Qwen3.8 Flash Next FP8
modelQwen3.8-Flash-Next is an experimental Qwen4-architecture preview: multimodal MoE, 125B parameters (6B active) plus 51B n-gram embeddings, native 262K context.
Qwen3.8 Flash Next FP4
modelQwen3.8-Flash-Next is an experimental Qwen4-architecture preview: a multimodal MoE with 125B parameters (6B active) plus 51B n-gram embeddings and native 262K context.
GLM-5.3
modelGLM-5.3 uses the same base model as GLM-5.2 — every gain comes from post-training. Compared with GLM-5.2, it is much better at complex coding and long-horizon tasks
Nemotron 3 Embed 8B
modelNVIDIA Nemotron-3-Embed-8B - an 8B-parameter text embedding model producing 4096-dimensional embeddings, with matryoshka support for truncating to smaller dimensions. Well suited for retrieval, semantic search, and RAG workloads.
Qwen3.8 27B
modelQwen3.8-27B: 27B-parameter vision-language model from Qwen with hybrid linear/full attention (4:1 interval) and 262K context.
DeepSeek-V4-Pro-0813
modelDeepSeek-V4-Pro-0813 is the official release of DeepSeek-V4-Pro, superseding the preview version, with greatly enhanced agentic capabilities and performance improvements that are especially pronounced in production environments. It is built on the DeepSeek-V4-Pro (Preview) model structure, with a DSpark speculative decoding module attached.
Qwen3.8-2.4T-A95B
modelQwen3.8-2.4T-A95B is Alibaba's most capable Qwen model to date, a 2.4T-parameter sparse MoE with ~95B active parameters. It is built for autonomous, long-horizon work: multi-day coding runs, research-paper reproduction and self-improvement.
Qwen 3.8 Max
modelBuilt on the architectural foundation of Qwen3.5, Qwen3.8 delivers substantial gains across coding, professional work, research, and long-horizon agentic tasks. Beyond answering harder questions, Qwen3.8 is designed to carry complex, multi-step tasks through to completion with greater reliability.
Muse Glimmer 30B
modelMuse Glimmer 30B is a dense causal language model distilled from Muse Spark and purpose-built for autonomous agentic work. It combines multi-step reasoning, reliable schema-based tool calling, and failure recovery with multimodal understanding via a ~1.8B ViT-G/14 perception encoder, supporting interleaved text and image input, a 131K+ context window, and selectable reasoning strength (low through xhigh). Trained on data from over 100 languages, Muse Glimmer performs strongly for its size class on agentic benchmarks including MCP Atlas, DeepSearch QA, Gaia2 and SWE-Bench Pro, and is released under Apache 2.0.
Nemotron Lightning 3.5 30B A3B
modelNemotron-Lightning-3.5-30B-A3B is a 30B-parameter Mixture-of-Experts language model (3B active) from NVIDIA's Nemotron-H family, built on a hybrid Mamba-Transformer architecture for efficient long-context inference. Like other models in the family, it responds to queries by first generating a reasoning trace and then concluding with a final response, with reasoning behavior configurable through a flag in the chat template. It includes a multi-token prediction (MTP) speculative decoding head for low-latency serving.
DeepSeek-V4-Flash-0731
modelDeepSeek-V4-Flash-0731 is the official release of DeepSeek-V4-Flash, superseding the preview version, with substantially enhanced agentic capabilities. It has the same model structure as DeepSeek-V4-Flash-DSpark, i.e. it comes with a speculative decoding module attached.
Inkling Small
modelA general-purpose multimodal model that accepts text, image, and audio inputs and generates text outputs.
Kimi K3
modelKimi K3 is Kimi’s most capable flagship model to date, with 2.8 trillion parameters. It is built on Kimi Delta Attention (KDA), a hybrid linear attention mechanism, and Attention Residuals, with native visual understanding and a 1M-token context window. It is the world’s first open-source model in the 3-trillion-parameter class, designed for frontier intelligence scenarios including long-horizon coding, knowledge work, and reasoning.
Ling 3 Flash
modelLing 3 Flash is currently planned as a 124B-parameter Mixture-of-Experts (MoE) model, with approximately 5.1B parameters activated per token. The model is designed with token efficiency and production-scale agentic inference as key priorities, enabling developers to complete more useful work within constrained token, latency, and serving-cost budgets.
Inkling
modelInkling - Thinking Machines multimodal (audio+vision) MoE. Inkling is the first open-weights model released by Thinking Machines Lab. It is a 975B Mixture-of-Experts with 41B active parameters, trained natively across text, image, and audio. Built as a broad generalist foundation for fine-tuning, with controllable thinking effort.
Qwen3.5 4B
modelBase model of 4B model from Qwen 3.5 series
GLM 5.2 FP8
modelGLM-5.2 FP8 checkpoint from zai-org/GLM-5.2-FP8
PaddleOCR VL 1.6
modelPaddleOCR-VL-1.6 is a compact ~0.9B vision-language model for document parsing (OCR, tables, formulas, charts, seals); SOTA on OmniDocBench v1.6. Supports image input, 131K context.
Voyage 4 Large
modelvoyage-4-large is Voyage AI's highest-quality, general-purpose (including multilingual) embedding model optimized for retrieval/search and AI applications. voyage-4-large supports embeddings in 2048, 1024, 512, and 256 dimensions, with multiple quantization options.
Page 1 of 13