Skip to main content

Cloudflare Workers AI Models

69 entities

@cf/cloudflare/clef

model

Clef is a 27B multimodal decision model that turns a state and a schema of typed questions into decisions. It reads the state as text, JSON, images, or video, and returns a probability for every allowed option of every question.

@cf/cloudflare/clef-flash

model

Clef-flash is a fast 9B multimodal decision model that turns a state and a schema of typed questions into decisions. It reads the state as text, JSON, images, or video, and returns a probability for every allowed option of every question.

@cf/swiss-ai/apertus-v1.5-8b

model

Apertus 1.5 is an 8B parameter language model designed to advance the state of multilingual, multimodal, fully open, and transparent AI. The models support a wide range of languages, handle contexts of up to 262,144 tokens, and it uses only fully open training data whilst delivering performance comparable to other models of similar size. For access, please fill out this form: https://forms.gle/kgvmkr6ucHN3xNuH6

@cf/utter-project/eurollm-9b-it

model

EuroLLM-9B is a 9B parameter model trained on 4 trillion tokens divided across the considered languages and several data sources: Web data, parallel data (en-xx and xx-en), and high-quality datasets. For access, please fill out this form: https://forms.gle/kgvmkr6ucHN3xNuH6

@cf/zai-org/glm-5.3

model

GLM-5.3 is Z.ai's flagship agentic coding model, pairing a 1M-token context window with reasoning, function calling, and structured outputs to power multi-step, tool-driven development workflows.

@cf/zai-org/glm-5.3-flash

model

The first natively multimodal model in the GLM-5 series. With 320B total parameters and just 18B active parameters, it outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic benchmarks.

@cf/qwen/qwen3.8-27b

model

Qwen 3.8 27B is a 27-billion-parameter instruction-tuned language model from Alibaba's Qwen family, designed for vision, efficient general-purpose text generation and agentic workloads.

@cf/deepseek-ai/deepseek-v4-flash-0731

model

DeepSeek-V4-Flash-0731 is the official release of DeepSeek-V4-Flash, superseding the preview version, with substantially enhanced agentic capabilities.

@cf/deepseek-ai/deepseek-v4-pro-0813

model

DeepSeek V4 Pro is a high-capability reasoning model from DeepSeek with a one million token context window, built for long-horizon agentic workflows and complex, multi-step problem-solving

@cf/moondream/moondream3.1-9B-A2B

model

Moondream 3 is a fast, efficient 9B mixture-of-experts vision language model (2B active parameters) that delivers frontier-level visual reasoning for tasks like object detection, pointing, OCR, and structured output.

@cf/zai-org/glm-5.2

model

Z.ai's flagship agentic coding model

@cf/moonshotai/kimi-k2.7-code

model

Kimi K2.7 is a frontier-scale open-source 1T parameter model with a 262.1k context window, multi-turn tool calling, vision inputs, and structured outputs for agentic workloads.

@cf/moonshotai/kimi-k2.6

model

Kimi K2.6 is a frontier-scale open-source 1T parameter model with a 262.1k context window, multi-turn tool calling, vision inputs, and structured outputs for agentic workloads.

@cf/google/gemma-4-26b-a4b-it

model

Gemma 4 is Google's most intelligent family of open models, built from Gemini 3 research to maximize intelligence-per-parameter.

@cf/nvidia/nemotron-3-120b-a12b

model

NVIDIA Nemotron 3 Super is a hybrid MoE model with leading accuracy for multi-agent applications and specialized agentic AI systems.

@cf/zai-org/glm-4.7-flash

model

GLM-4.7-Flash is a fast and efficient multilingual text generation model with a 131,072 token context window. Optimized for dialogue, instruction-following, and multi-turn tool calling across 100+ languages.

@cf/ai4bharat/indictrans2-en-indic-1B

model

IndicTrans2 is the first open-source transformer-based multilingual NMT model that supports high-quality translations across all the 22 scheduled Indic languages

@cf/aisingapore/gemma-sea-lion-v4-27b-it

model

SEA-LION stands for Southeast Asian Languages In One Network, which is a collection of Large Language Models (LLMs) which have been pretrained and instruct-tuned for the Southeast Asia (SEA) region.

@cf/baai/bge-base-en-v1.5

model

BAAI general embedding (Base) model that transforms any given text into a 768-dimensional vector

@cf/baai/bge-large-en-v1.5

model

BAAI general embedding (Large) model that transforms any given text into a 1024-dimensional vector

@cf/baai/bge-m3

model

Multi-Functionality, Multi-Linguality, and Multi-Granularity embeddings model.

@cf/baai/bge-reranker-base

model

Different from embedding model, reranker uses question and document as input and directly output similarity instead of embedding. You can get a relevance score by inputting query and passage to the reranker. And the score can be mapped to a float value in [0,1] by sigmoid function.

@cf/baai/bge-small-en-v1.5

model

BAAI general embedding (Small) model that transforms any given text into a 384-dimensional vector

@cf/black-forest-labs/flux-1-schnell

model

FLUX.1 [schnell] is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions.

@cf/black-forest-labs/flux-2-dev

model

FLUX.2 [dev] is an image model from Black Forest Labs where you can generate highly realistic and detailed images, with multi-reference support.

Page 1 of 3