⏰ OpenRouter Limited-Time Free Models
| Model | Expires | Context | Input/Output Price |
|---|---|---|---|
| NVIDIA: Llama 3.3 Nemotron Super 49B V1.5 | Jul 17 | 131K | $0.40/M tokens |
| Meta: Llama 3.2 11B Vision Instruct | Jul 17 | 131K | $0.35/M tokens |
| Qwen: Qwen3 Next 80B A3B Instruct (free) | Jul 19 | 262K | Free |
| Qwen: Qwen3 Coder 480B A35B (free) | Jul 19 | 1M | Free |
| Venice: Uncensored (free) | Jul 19 | 33K | Free |
| Meta: Llama 3.3 70B Instruct (free) | Jul 19 | 131K | Free |
| Meta: Llama 3.2 3B Instruct (free) | Jul 19 | 131K | Free |
| Nous: Hermes 3 405B Instruct (free) | Jul 19 | 131K | Free |
| Tencent: Hy3 (free) | Jul 21 | 262K | Free |
| OpenAI: GPT-5.3 Chat | Aug 10 | 128K | $1.75/M input, $14/M output |
Free models will expire Jul 19-21. The GPT-5.3 Chat free tier (upgraded from GPT-5.2) runs until Aug 10 at reduced pricing.
🤯 Major Model Releases This Week
Anthropic: Claude Sonnet 5, Opus 4.8 & Fable 5
Anthropic has unveiled three new Claude models on OpenRouter: Claude Sonnet 5 ($2/$10 per M tokens), Claude Opus 4.8 ($5/$25 per M, with a fast variant at $10/$50), and Claude Fable 5 ($10/$50 per M). All three feature 1M context windows. This is Anthropic’s largest simultaneous model launch, spanning the full spectrum from cost-efficient to frontier intelligence.
OpenAI: GPT-5.6 Luna/Terra/Sol Family Now Available
OpenAI’s GPT-5.6 series has gone live on OpenRouter with a three-tier pricing structure:
- GPT-5.6 Luna ($1/$6 per M, 1.05M ctx) — fastest, most cost-effective tier
- GPT-5.6 Terra ($2.50/$15 per M, 1.05M ctx) — balanced performance
- GPT-5.6 Sol ($5/$30 per M, 1.05M ctx) — flagship reasoning tier
Pro variants of each tier offer enhanced reliability at the same pricing. Additionally, GPT-5.3 Chat ($1.75/$14 per M, 128K ctx) replaces GPT-5.2 Chat, and GPT-5.3 Codex ($1.75/$14 per M, 400K ctx) debuts as a code-specialized variant.
Qwen: Entire 3.6 Family & 3.7 Plus/Max Released
Alibaba’s Qwen team has dropped a massive model wave:
- Qwen 3.7 Plus ($0.32/$1.28 per M, 1M ctx)
- Qwen 3.7 Max ($1.25/$3.75 per M, 1M ctx)
- Qwen 3.6 Flash ($0.19/$1.12 per M, 1M ctx) — ultra-fast inference
- Qwen 3.6 35B A3B ($0.14/$1 per M, 262K ctx) — highly efficient MoE
- Qwen 3.6 27B ($0.45/$2.70 per M, 262K ctx)
- Qwen 3.6 Plus ($0.33/$1.95 per M, 1M ctx)
- Qwen 3.6 Max Preview ($1.04/$6.24 per M, 262K ctx)
Qwen 3.6 Flash at $0.19/M tokens makes it one of the most affordable capable models on the market.
Google: Gemini 3.5 Flash
Google has released Gemini 3.5 Flash ($1.50/$9 per M tokens, 1M context) — a faster iteration of the Gemini 3.x Flash line, optimized for high-throughput applications.
⭐ GitHub Trending: AI Edition
- AlephAITech/WorkBuddyGuide — Open-source guide to mastering WorkBuddy through real-world workflows. ⭐ 757
- William-Lu-stack/Flawless — AI SRE AgenticOps for Kubernetes and cloud infrastructure. ⭐ 619
- yetone/kill-ai-slop — A field guide to the visual & copy tics of AI-generated products — and an Agent Skill that scans you. ⭐ 451
- QuantumByteOSS/quantumbyte — Open-source app builder engine — intent to working app. ⭐ 307
- dongguatanglinux/grok-build-auth — Protocol research client for x.ai xAI SSO → Grok Build OAuth → CLI tooling. ⭐ 228
🤗 Hugging Face New Uploads
Notable recent uploads include whisper-large-v3-turbo Thai-English adapters, GGUF quantized Gemma-4-12b variants, and several new safetensors models covering Catalan language models, SWE-bench specialized models, and checkpoint collections.
💡 Key Trends
-
The Model Release Avalanche — July 15 alone saw Anthropic, OpenAI, Google, and Alibaba all ship major new models within hours of each other. The market is experiencing an unprecedented cadence of frontier model releases as labs compete for developer mindshare on OpenRouter. The sheer volume — Claude Sonnet 5, Opus 4.8, Fable 5, GPT-5.6 x3 tiers, GPT-5.3 x2 variants, Qwen 3.7 x2 + 3.6 x5, Gemini 3.5 Flash — suggests we’ve entered a new phase of hyperscale model proliferation.
-
Price War at the Low End — Qwen 3.6 Flash at $0.19/M and the 35B A3B MoE at $0.14/M represent the bleeding edge of cost-efficient inference. Combined with free-tier access to Qwen3 Coder (1M ctx) and Llama 3.3 70B through Jul 19, the barrier to running capable models has never been lower.
-
Anthropic’s Multi-Model Strategy — Launching Sonnet 5, Opus 4.8, and Fable 5 simultaneously signals Anthropic’s intent to cover every price-performance tier rather than offering a single flagship. This mirrors OpenAI’s GPT-5.6 Luna/Terra/Sol tiering and suggests the industry is converging on model families rather than individual releases.
-
Free Tier Window Closing — Multiple free models expire Jul 17-19 (NVIDIA Nemotron, Meta Llama Vision, and the Qwen/Meta/Nous free tier block). Users should plan migrations or budget for paid usage.
Today’s Snapshot
- OpenRouter limited-free models: 10
- Hugging Face trending models: 10 new uploads
- GitHub trending AI repos: 5
- Major model releases this week: 15+ (Anthropic 3, OpenAI 5, Qwen 7, Google 1, NVIDIA Nemotron 3 free variants)
- Top story: Anthropic Claude Sonnet 5 + OpenAI GPT-5.6 series launch on OpenRouter
Covering Jul 13–15, 2026. Sources: OpenRouter, Hugging Face, GitHub.