Skip to main content

Fireworks AI Models

308 entities

MiMo-V2.6-Pro-RL

model

MiMo-V2.6-Pro-RL is a 1T-parameter sparse MoE omnimodal model (text, image, video, audio) with 1M-token context, designed as Xiaomi’s flagship agentic model. It scales a unified reinforcement-learning pipeline across coding, general agents, visual tasks, and cybersecurity using groupwise agentic grading and speculative decoding to drive self-improvement and strong benchmark performance

Ember-1

model

Ember-1 is a specialized model from Fireworks. Built on Kimi K3, it produces shorter reasoning traces, using approximately 40% fewer tokens while maintaining comparable quality across our evaluations

DeepSeek V4.1 Flash

model

DeepSeek-V4.1-Flash is a multimodal Mixture-of-Experts model with 552B backbone parameters that natively processes images and text at up to 1M token context. Its Causal Encoder-Decoder architecture activates only 8B parameters during prefill and 16B during decode, cutting the KV cache footprint to roughly a quarter of DeepSeek-V4-Flash for cost-efficient agentic workloads.

Ling 3 Flash Fin

model

Ling-3.0-flash-Fin is the first finance-enhanced model in the Ant Ling family. Developed by Ant Group with leading financial institutions and domain experts, it extends Ling-3.0-flash through continued training on high-quality financial data. With 124B total parameters, 5.1B activated parameters, and a 256K context window, the model combines financial expertise with efficient inference for long-horizon agent workflows.

DeepSeek-V4-Flash-Vision-Exp

model

DeepSeek-V4-Flash-Vision-Exp is the first experimental multimodal model in the DeepSeek-V4 family. It builds on DeepSeek-V4-Flash by adding visual modules and continued training for visual understanding, with substantially improved multimodal agent capabilities while remaining comparable on text-only agent tasks. It has the same model structure as DeepSeek-V4-Flash-DSpark, i.e. it comes with a speculative decoding module attached.

GLM 5.3 Flash

model

GLM-5.3-Flash is the first natively multimodal model in the GLM-5 series with 320B total parameters and 18B active parameters. It incorporates several architectural improvements over GLM-5.2 including a hybrid architecture, sharply reducing long-context serving costs while preserving precise long-context and Manifold-Constrained Hyper-Connections (mHC) to further improve scaling efficiency.

Qwen3.8 Flash Next FP8

model

Qwen3.8-Flash-Next is an experimental Qwen4-architecture preview: multimodal MoE, 125B parameters (6B active) plus 51B n-gram embeddings, native 262K context.

Qwen3.8 Flash Next FP4

model

Qwen3.8-Flash-Next is an experimental Qwen4-architecture preview: a multimodal MoE with 125B parameters (6B active) plus 51B n-gram embeddings and native 262K context.

GLM-5.3

model

GLM-5.3 uses the same base model as GLM-5.2 — every gain comes from post-training. Compared with GLM-5.2, it is much better at complex coding and long-horizon tasks

Nemotron 3 Embed 8B

model

NVIDIA Nemotron-3-Embed-8B - an 8B-parameter text embedding model producing 4096-dimensional embeddings, with matryoshka support for truncating to smaller dimensions. Well suited for retrieval, semantic search, and RAG workloads.

Qwen3.8 27B

model

Qwen3.8-27B: 27B-parameter vision-language model from Qwen with hybrid linear/full attention (4:1 interval) and 262K context.

DeepSeek-V4-Pro-0813

model

DeepSeek-V4-Pro-0813 is the official release of DeepSeek-V4-Pro, superseding the preview version, with greatly enhanced agentic capabilities and performance improvements that are especially pronounced in production environments. It is built on the DeepSeek-V4-Pro (Preview) model structure, with a DSpark speculative decoding module attached.

Qwen3.8-2.4T-A95B

model

Qwen3.8-2.4T-A95B is Alibaba's most capable Qwen model to date, a 2.4T-parameter sparse MoE with ~95B active parameters. It is built for autonomous, long-horizon work: multi-day coding runs, research-paper reproduction and self-improvement.

Qwen 3.8 Max

model

Built on the architectural foundation of Qwen3.5, Qwen3.8 delivers substantial gains across coding, professional work, research, and long-horizon agentic tasks. Beyond answering harder questions, Qwen3.8 is designed to carry complex, multi-step tasks through to completion with greater reliability.

Muse Glimmer 30B

model

Muse Glimmer 30B is a dense causal language model distilled from Muse Spark and purpose-built for autonomous agentic work. It combines multi-step reasoning, reliable schema-based tool calling, and failure recovery with multimodal understanding via a ~1.8B ViT-G/14 perception encoder, supporting interleaved text and image input, a 131K+ context window, and selectable reasoning strength (low through xhigh). Trained on data from over 100 languages, Muse Glimmer performs strongly for its size class on agentic benchmarks including MCP Atlas, DeepSearch QA, Gaia2 and SWE-Bench Pro, and is released under Apache 2.0.

Nemotron Lightning 3.5 30B A3B

model

Nemotron-Lightning-3.5-30B-A3B is a 30B-parameter Mixture-of-Experts language model (3B active) from NVIDIA's Nemotron-H family, built on a hybrid Mamba-Transformer architecture for efficient long-context inference. Like other models in the family, it responds to queries by first generating a reasoning trace and then concluding with a final response, with reasoning behavior configurable through a flag in the chat template. It includes a multi-token prediction (MTP) speculative decoding head for low-latency serving.

DeepSeek-V4-Flash-0731

model

DeepSeek-V4-Flash-0731 is the official release of DeepSeek-V4-Flash, superseding the preview version, with substantially enhanced agentic capabilities. It has the same model structure as DeepSeek-V4-Flash-DSpark, i.e. it comes with a speculative decoding module attached.

Inkling Small

model

A general-purpose multimodal model that accepts text, image, and audio inputs and generates text outputs.

Kimi K3

model

Kimi K3 is Kimi’s most capable flagship model to date, with 2.8 trillion parameters. It is built on Kimi Delta Attention (KDA), a hybrid linear attention mechanism, and Attention Residuals, with native visual understanding and a 1M-token context window. It is the world’s first open-source model in the 3-trillion-parameter class, designed for frontier intelligence scenarios including long-horizon coding, knowledge work, and reasoning.

Ling 3 Flash

model

Ling 3 Flash is currently planned as a 124B-parameter Mixture-of-Experts (MoE) model, with approximately 5.1B parameters activated per token. The model is designed with token efficiency and production-scale agentic inference as key priorities, enabling developers to complete more useful work within constrained token, latency, and serving-cost budgets.

Inkling

model

Inkling - Thinking Machines multimodal (audio+vision) MoE. Inkling is the first open-weights model released by Thinking Machines Lab. It is a 975B Mixture-of-Experts with 41B active parameters, trained natively across text, image, and audio. Built as a broad generalist foundation for fine-tuning, with controllable thinking effort.

Qwen3.5 4B

model

Base model of 4B model from Qwen 3.5 series

GLM 5.2 FP8

model

GLM-5.2 FP8 checkpoint from zai-org/GLM-5.2-FP8

PaddleOCR VL 1.6

model

PaddleOCR-VL-1.6 is a compact ~0.9B vision-language model for document parsing (OCR, tables, formulas, charts, seals); SOTA on OmniDocBench v1.6. Supports image input, 131K context.

Voyage 4 Large

model

voyage-4-large is Voyage AI's highest-quality, general-purpose (including multilingual) embedding model optimized for retrieval/search and AI applications. voyage-4-large supports embeddings in 2048, 1024, 512, and 256 dimensions, with multiple quantization options.

Page 1 of 13