⏰ OpenRouter Limited-Time Free Models

Model Context Pricing Expires
Arcee AI: Trinity Mini 131K tokens $0.045/M input → $0.15/M output Today — Jul 9 (last day!)
Tencent: Hy3 (Free) 262K tokens Free Jul 21 (12 days left)

The Poolside Laguna XS.2 and Google Gemini 2.5 Flash Lite free tiers have now expired. Arcee AI’s Trinity Mini — a compact 131K-context model — expires today, while Tencent Hy3 fresh from its recent debut still offers free inference through July 21.

🆕 New on OpenRouter

xAI: Grok 4.5

xAI’s latest frontier model Grok 4.5 has landed on OpenRouter with a massive 500,000-token context window — the largest among frontier-class models on the platform. Priced at $2.00/M tokens for input and $8.00/M for output, it positions itself as a premium reasoning and coding powerhouse alongside a “Grok Latest” alias tracking the newest release.

AionLabs: Aion-3.0 Family

AionLabs has debuted two new models: Aion-3.0 (full) and Aion-3.0-Mini. Both offer 131K-token context windows, with the full model priced at $3.00/M input and the Mini at a modest $0.70/M input — a competitive entry in the mid-tier reasoning segment.

Tencent Hy3

Tencent’s Hy3 (covered on Jul 7) continues to make waves with its 262K context at $0.14/M input — one of the most cost-effective frontier-class models available. The free tier is available through July 21.

  • csukuangfj2/sherpa-onnx-libs — Cross-platform ONNX-based speech recognition libraries, building on the well-known sherpa-onnx project for offline ASR and TTS. (⭐ 4)
  • mradermacher/Mistral-Heretica-12B-i1-GGUF — A GGUF-quantized Mistral variant with uncensored capabilities, optimized for local deployment via llama.cpp. (imatrix, conversational)
  • mradermacher/LM-Lexicon-8B-Dense-Oxford-GGUF — An 8B-parameter “dense” model integrating lexical/Oxford reference knowledge, available in GGUF format for efficient local inference.
  • mradermacher/Katarau-9B-ru-RP-nsfw-GGUF — A 9B Russian-language roleplay model in GGUF format, designed for uncensored conversational and creative writing tasks in Russian.
  • lsxi77777/Wat3R — A 3D reconstruction model (safetensors, PyTorch) using a model hub mixin architecture, likely related to neural 3D scene understanding.
  • gghfez/c4ai-command-r-v01-jacobian-lens — An interpretability release applying Jacobian Lens to Cohere’s Command-R model, offering researchers a window into internal model representations. (base_model: CohereLabs/c4ai-command-r-v01)
  • dodoelreedy/medical-chatbot — An MIT-licensed medical chatbot model, targeting healthcare Q&A and patient interaction use cases.
  • darsh23262/skin-detection-yolo26 — A YOLOv26-based skin detection model (MIT license), extending the YOLO object detection lineage to medical/dermatological imaging.
Repo Description Stars
synthetic-sciences/openscience Open-source AI workbench for scientific research — unified search, coding, data analysis ⭐ 1,864 (+781 in 2 days)
jamesob/local-llm Comprehensive guide to running LLMs locally — Shell-based setup from scratch ⭐ 1,290 (+173 in 2 days)
simonlin1212/Vibe-Research Personal trading research agent for A-shares, US stocks, and HK stocks — daily review, news radar, portfolio tracking ⭐ 575
ai4s-research/open-science Open Science Desktop — local-first, model-agnostic AI research workbench for macOS & Windows ⭐ 456
Pluviobyte/rnskill A collection of AI Agent Skills, focused on automating developer workflows ⭐ 387
  1. xAI’s Grok 4.5 Goes Public on OpenRouter. With a 500K-token context window and pricing that puts it alongside premium frontier models, Grok 4.5 marks xAI’s most aggressive push yet into the model-as-a-service market. The inclusion of a “Grok Latest” alias suggests rapid iteration — users who point to the alias will automatically get future improvements without changing endpoints.

  2. Meta’s Hardware Innovation: Reusing Old RAM. Meta is building custom bridge chips that let them reuse older-generation RAM in new server builds — a significant infrastructure cost-saver at hyperscale. Instead of discarding perfectly functional DRAM when upgrading CPU platforms, the bridge chip adapts older memory modules to newer memory controllers. This is a rare glimpse of the hardware-level cost optimization that powers Meta’s massive AI training clusters.

  3. AionLabs Enters the Arena. AionLabs’ dual release of Aion-3.0 and Aion-3.0-Mini adds a new player to the increasingly crowded mid-tier reasoning model space. With 131K context and competitive pricing, they’re targeting developers who need capable models without frontier-tier costs.

  4. OpenScience and the Scientific AI Workbench Trend. synthetic-sciences/openscience surged 781 stars in just two days (from 1,083 to 1,864), reflecting mounting interest in AI-powered research platforms. The rise of local-first, model-agnostic workbenches (also seen with ai4s-research/open-science) suggests a growing demand for tools that unify literature search, code execution, and data analysis in a single AI-native environment.

  5. GGUF Ecosystem Continues to Thrive. With multiple mradermacher quantized releases (Mistral-Heretica, LM-Lexicon, Katarau), the GGUF format remains the dominant standard for local LLM deployment. The diversity — Russian roleplay, uncensored variants, lexical knowledge models — shows how quantization democratizes access to specialized models.

Today’s Snapshot

  • 10 new Hugging Face models (speech recognition, GGUF quants, Jacobian lens interpretability, 3D reconstruction, medical chatbot, skin detection)
  • 6 new OpenRouter models (Grok 4.5, Aion-3.0 family, Tencent Hy3)
  • 5 trending GitHub repos (scientific AI workbench, local LLM guide, trading agent, AI Agent skills)
  • 0 new papers found
  • 2 limited-time free models (Trinity Mini — ends today, Hy3 free — ends Jul 21)