Cloudflare Workers AI Models
69 entities
@cf/cloudflare/clef
modelClef is a 27B multimodal decision model that turns a state and a schema of typed questions into decisions. It reads the state as text, JSON, images, or video, and returns a probability for every allowed option of every question.
@cf/cloudflare/clef-flash
modelClef-flash is a fast 9B multimodal decision model that turns a state and a schema of typed questions into decisions. It reads the state as text, JSON, images, or video, and returns a probability for every allowed option of every question.
@cf/swiss-ai/apertus-v1.5-8b
modelApertus 1.5 is an 8B parameter language model designed to advance the state of multilingual, multimodal, fully open, and transparent AI. The models support a wide range of languages, handle contexts of up to 262,144 tokens, and it uses only fully open training data whilst delivering performance comparable to other models of similar size. For access, please fill out this form: https://forms.gle/kgvmkr6ucHN3xNuH6
@cf/utter-project/eurollm-9b-it
modelEuroLLM-9B is a 9B parameter model trained on 4 trillion tokens divided across the considered languages and several data sources: Web data, parallel data (en-xx and xx-en), and high-quality datasets. For access, please fill out this form: https://forms.gle/kgvmkr6ucHN3xNuH6
@cf/zai-org/glm-5.3
modelGLM-5.3 is Z.ai's flagship agentic coding model, pairing a 1M-token context window with reasoning, function calling, and structured outputs to power multi-step, tool-driven development workflows.
@cf/zai-org/glm-5.3-flash
modelThe first natively multimodal model in the GLM-5 series. With 320B total parameters and just 18B active parameters, it outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic benchmarks.
@cf/qwen/qwen3.8-27b
modelQwen 3.8 27B is a 27-billion-parameter instruction-tuned language model from Alibaba's Qwen family, designed for vision, efficient general-purpose text generation and agentic workloads.
@cf/deepseek-ai/deepseek-v4-flash-0731
modelDeepSeek-V4-Flash-0731 is the official release of DeepSeek-V4-Flash, superseding the preview version, with substantially enhanced agentic capabilities.
@cf/deepseek-ai/deepseek-v4-pro-0813
modelDeepSeek V4 Pro is a high-capability reasoning model from DeepSeek with a one million token context window, built for long-horizon agentic workflows and complex, multi-step problem-solving
@cf/moondream/moondream3.1-9B-A2B
modelMoondream 3 is a fast, efficient 9B mixture-of-experts vision language model (2B active parameters) that delivers frontier-level visual reasoning for tasks like object detection, pointing, OCR, and structured output.
@cf/zai-org/glm-5.2
modelZ.ai's flagship agentic coding model
@cf/moonshotai/kimi-k2.7-code
modelKimi K2.7 is a frontier-scale open-source 1T parameter model with a 262.1k context window, multi-turn tool calling, vision inputs, and structured outputs for agentic workloads.
@cf/moonshotai/kimi-k2.6
modelKimi K2.6 is a frontier-scale open-source 1T parameter model with a 262.1k context window, multi-turn tool calling, vision inputs, and structured outputs for agentic workloads.
@cf/google/gemma-4-26b-a4b-it
modelGemma 4 is Google's most intelligent family of open models, built from Gemini 3 research to maximize intelligence-per-parameter.
@cf/nvidia/nemotron-3-120b-a12b
modelNVIDIA Nemotron 3 Super is a hybrid MoE model with leading accuracy for multi-agent applications and specialized agentic AI systems.
@cf/zai-org/glm-4.7-flash
modelGLM-4.7-Flash is a fast and efficient multilingual text generation model with a 131,072 token context window. Optimized for dialogue, instruction-following, and multi-turn tool calling across 100+ languages.
@cf/ai4bharat/indictrans2-en-indic-1B
modelIndicTrans2 is the first open-source transformer-based multilingual NMT model that supports high-quality translations across all the 22 scheduled Indic languages
@cf/aisingapore/gemma-sea-lion-v4-27b-it
modelSEA-LION stands for Southeast Asian Languages In One Network, which is a collection of Large Language Models (LLMs) which have been pretrained and instruct-tuned for the Southeast Asia (SEA) region.
@cf/baai/bge-base-en-v1.5
modelBAAI general embedding (Base) model that transforms any given text into a 768-dimensional vector
@cf/baai/bge-large-en-v1.5
modelBAAI general embedding (Large) model that transforms any given text into a 1024-dimensional vector
@cf/baai/bge-m3
modelMulti-Functionality, Multi-Linguality, and Multi-Granularity embeddings model.
@cf/baai/bge-reranker-base
modelDifferent from embedding model, reranker uses question and document as input and directly output similarity instead of embedding. You can get a relevance score by inputting query and passage to the reranker. And the score can be mapped to a float value in [0,1] by sigmoid function.
@cf/baai/bge-small-en-v1.5
modelBAAI general embedding (Small) model that transforms any given text into a 384-dimensional vector
@cf/black-forest-labs/flux-1-schnell
modelFLUX.1 [schnell] is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions.
@cf/black-forest-labs/flux-2-dev
modelFLUX.2 [dev] is an image model from Black Forest Labs where you can generate highly realistic and detailed images, with multi-reference support.
Page 1 of 3