210 results
Foundation models you can call or deploy
by EleutherAI·Jun 2021
Open-source 6-billion parameter autoregressive language model by EleutherAI, trained on The Pile; released with weights as an alternative to GPT-3.
by Meta·Dec 2023
Meta's open-weight safety classifier for filtering harmful LLM inputs/outputs; initial release was a Llama 2 7B instruction-tuned model, followed by v2, v3, and v4.
by Mistral AI·Jul 2025
Mistral's first open audio-input model family (Mini 3B and Small 24B), Apache 2.0; multilingual speech understanding, transcription, summarization and Q&A over audio.
by Cohere·Aug 2026
A 2.3-billion-parameter multimodal document parser that converts complex documents into structured Markdown with tables, forms, images, and bounding boxes.
by Z.ai·Aug 2026
A natively multimodal mixture-of-experts model with 320 billion total parameters and 18 billion active parameters, available through open weights and an API.
by Google·Jun 2025
Fast, cost-efficient Gemini 2.5 model with hybrid reasoning that lets developers toggle thinking on or off and set thinking budgets per request.
by Z.ai·Sep 2025
357B MoE with 200K context (up from 128K in 4.5), MIT license; improved coding, tool-use during inference, and agentic benchmark performance.
by OpenAI·Jun 2025
Higher-compute variant of OpenAI's o3 reasoning model, available to ChatGPT Pro users and via API with tool access (web, files, Python) for high-reliability tasks.
by Moonshot AI·Nov 2025
1T/32B-active MoE with 256K context, native INT4 QAT, Modified MIT license; deep-thinking variant with 200-300 sequential tool calls for agentic tasks.
by Mistral AI·Jul 2024
12B dense LLM co-developed with NVIDIA, 128k context, Apache 2.0; multilingual and quantization-aware, positioned as a drop-in replacement for Mistral 7B.
by Anthropic·Sep 2026
Anthropic's generally available model for long-running coding, agentic, research, and knowledge-work tasks, with safeguards for high-risk cybersecurity and biology requests.
by Google·Aug 2026
A speech-transcription model supporting more than 85 languages, speaker diarization, word-level timestamps, custom vocabulary, and live or file-based audio.
by Tencent·Aug 2026
A 770-billion-parameter mixture-of-experts language model with 49 billion active parameters and a context window exceeding one million tokens, released with open weights and API access.
by Microsoft·May 2024
Microsoft's 4.2B multimodal variant of Phi-3 that accepts text + image input, small enough to run on-device for OCR, chart understanding, and vision-augmented chat.
by OpenAI·Mar 2026
Smaller cost-optimized variant of GPT-5.4, available on the free tier via the Thinking feature and as an API model; ~4x pricier than GPT-5 equivalents in the OpenAI API.
by Google·Feb 2026
Incremental Gemini 3 Pro upgrade with stronger core reasoning, verified 77.1% on ARC-AGI-2, shipped in preview across AI Studio, Antigravity, and Vertex AI.
by Moonshot AI·Oct 2025
48B/3B-active hybrid linear-attention MoE (Kimi Delta Attention + MLA at 3:1), 1M context, MIT license; 6x faster decode via ~75% KV cache reduction.
Microsoft's 7B dense Phi-3 SLM with 8K and 128K context variants (MIT license), benchmarked as a mid-tier textbook-trained model between Phi-3 mini and medium.
Anthropic's access-controlled model for advanced cybersecurity and biology research, available to vetted organizations through trusted-access programs.
A generally available multimodal model for video generation, extension, interpolation, and editing at resolutions up to 4K.
by Alibaba·Aug 2026
An open-weight multimodal model and experimental preview of the architecture planned for Qwen4, designed around sparse attention and efficient long-context inference.
by Google·Dec 2024
Fast, low-cost Gemini 2.0 model with native multimodal output (images plus steerable multilingual text-to-speech) and built-in tool use including Search.
by Moonshot AI·Jul 2025
1T/32B-active MoE with 256K context and MLA attention, Modified MIT license; 384-expert (8/token) agentic-focused foundation model from Moonshot.
A generally available multimodal Gemini model for coding and agentic workflows with a one-million-token input context window.