196 results
Foundation models you can call or deploy
by OpenAI·Jun 2018
First Generative Pre-trained Transformer from OpenAI (2018), a 117M-parameter decoder-only model pre-trained on unlabeled text then fine-tuned for NLP tasks.
by EleutherAI·Jun 2021
Open-source 6-billion parameter autoregressive language model by EleutherAI, trained on The Pile; released with weights as an alternative to GPT-3.
by Anthropic·May 2025
Anthropic's Claude Opus 4, initial Claude 4 flagship for coding, reasoning, and agent workflows; retired June 2026 in favor of Opus 4.8.
by Google·May 2026
Gemini 3.5 Flash: fast agentic and coding model built for long-horizon tasks, rivaling flagship-model quality at Flash-tier latency and cost.
by xAI·Aug 2024
xAI's second-generation LLM, launched in beta on X with Grok-2 and Grok-2 mini; weights later open-sourced in 2025, using dense attention with MoE elements.
by Meta·Apr 2025
Meta's Llama 4 Maverick MoE model with 17B active / 128 experts (400B total), 1M token context, natively multimodal, under the Llama 4 Community License.
by Anthropic·Mar 2024
Smallest, fastest model in Anthropic's Claude 3 family for high-volume tasks; supports vision input and the Messages API. Retired April 2026.
by Mistral AI·Jul 2025
Mistral's first open audio-input model family (Mini 3B and Small 24B), Apache 2.0; multilingual speech understanding, transcription, summarization and Q&A over audio.
by OpenAI·Aug 2025
OpenAI's open-weight 120B-parameter reasoning model released under Apache 2.0, aimed at production self-hosting with function calling and structured outputs support.
by Google·Jun 2025
Smallest, cheapest Gemini 2.5 model, tuned for high-volume translation and classification with a 1M-token context and toggleable thinking budgets.
by Google·Feb 2026
Incremental Gemini 3 Pro upgrade with stronger core reasoning, verified 77.1% on ARC-AGI-2, shipped in preview across AI Studio, Antigravity, and Vertex AI.
by Midjourney·Apr 2025
Midjourney's V7 text-to-image model with improved aesthetics, prompt adherence, and image detail; accessible via Discord and the Midjourney web app.
by Google·Dec 2023
Smallest Gemini model, designed for on-device inference in Android (starting on Pixel 8 Pro) and Chrome, enabling local AI features without cloud calls.
by Alibaba·Jun 2025
Unified multimodal understanding-and-generation preview from Alibaba's Qwen team; text-to-image and instruction-based image editing accessible via Qwen Chat.
by Mistral AI·May 2025
Mistral's 24B open-weight agentic coding model, Apache 2.0, tuned for tool use in software-engineering tasks; built with All Hands AI, tops open-model SWE-Bench.
by Google·Dec 2024
Google's second-generation multimodal model family, built for the agentic era with native tool use, multimodal input, and image plus audio output.
by OpenAI·Feb 2019
OpenAI's 2019 transformer language model (up to 1.5B parameters), known for coherent zero-shot text generation; initially released in staged increments over misuse concerns.
by Mistral AI·Dec 2023
Mistral's sparse Mixture-of-Experts family; Mixtral 8x7B (46.7B total, ~12.9B active) under Apache 2.0, outperforming Llama 2 70B and GPT-3.5 on many benchmarks.
by Microsoft·May 2024
Microsoft's 7B dense Phi-3 SLM with 8K and 128K context variants (MIT license), benchmarked as a mid-tier textbook-trained model between Phi-3 mini and medium.
Microsoft's 4.2B multimodal variant of Phi-3 that accepts text + image input, small enough to run on-device for OCR, chart understanding, and vision-augmented chat.
by xAI·Jul 2025
Multi-agent high-compute variant of Grok 4 delivered via the SuperGrok Heavy tier; runs parallel agents on the same query for advanced reasoning benchmarks.
by xAI·Aug 2025
xAI's speedy, economical reasoning model built from scratch for agentic coding; optimized for tool use inside Cursor, GitHub Copilot, Cline, Roo Code, opencode, and Windsurf.
by Anthropic·Jul 2023
Anthropic's second-generation Claude model with a 100k-token context window for chat, summarization, and coding; retired from the API in July 2025.
by Moonshot AI·Nov 2025
1T/32B-active MoE with 256K context, native INT4 QAT, Modified MIT license; deep-thinking variant with 200-300 sequential tool calls for agentic tasks.