Models 32 events
Category RSS ↗Today
-
Limited sourcesModels
AA is the reason for Qwen3.8 27B shipped with xhigh
I know why Qwen3.8 27B shipped with xhigh reasoning as default, it's to do its best in benchmarks. Models from top labs often get benchmarked at multiple reasoning levels, but that same treatment doe…
Reddit r/LocalLLaMA -
Limited sourcesModelsAgents & Dev
Optimizing Qwen3.6 / Qwen3.8-27B on 16GB VRAM: Complete Benchmark Results and Setup Guide (~30-50tps at 32k to 72k context)
This post was made with AI. I tried to remove as much slop as possible and keep it straight to the point to save your time as I know how annoying AI slop posts can be, but I still wanted to retain al…
Reddit r/LocalLLaMA Yesterday
-
Limited sourcesModelsResearch
Benchmarked Qwen3.8-27B on 4x RTX 3090
A while back I made a post about my 4x3090 rig in a Silverstone RV-02 . Check it out if you're a conoissuer of OG PC cases. With the incredible Qwen 3.8 27B release I ran benchmarks. So in case you a…
Reddit r/LocalLLaMA -
Limited sourcesModels
Qwen3.8-27B Uncensored Aggressive is out with K_P quants and HauhauCS FastMTP (up to 3.02x TG)!
The dense Qwen release is back! Qwen3.8-27B Uncensored Aggressive is out with the complete K_P quant range, Vision, native NextN, and HauhauCS FastMTP. Aggressive here means no refusals, no personali…
Reddit r/LocalLLaMA -
Limited sourcesModelsBusiness
GPT-5.6 Sol Pricing Cut by 50%
openai / gpt-5.6-sol GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and mu…
Hacker News Frontpage -
Limited sourcesModels
Weirdly, no one talks about Temperature setting for the Qwen3.8 27b
Mind you, it is 1.0 by default, yet everyone is focused on how much the new model thinks, restricting the reasoning budget and/or dropping the reasoning level. Set the temperature to 0.7 and the mode…
Reddit r/LocalLLaMA -
Limited sourcesModels
I pushed Qwen3.8-27B to 99 tps single request and 1150 tps with a batch request on a RTX 3090
I'm back. Yesterday I released the first version of hyper-optimized Qwen3.8-27B inference engine for a RTX 3090, reaching 82 tps on single request and 672 peak. Over the last 24 hours I've been explo…
Reddit r/LocalLLaMA -
Limited sourcesModels
Qwen3.8 27B's result on Artificial Analysis is insane!
submitted by /u/FormOne2615 [link] [comments]
Reddit r/LocalLLaMA -
Limited sourcesModelsAgents & Dev
Qwen3.8 27B > Opus 5 Medium on Artificial Analysis Agentic Index
https://preview.redd.it/xh1rloaf4zjh1.png?width=1628&format=png&auto=webp&s=a536fae1b50b327f2bc55d4f4f874f94ae66e867 Thanks Qwen team! submitted by /u/secopsml [link] [comments]
Reddit r/LocalLLaMA -
Limited sourcesModels
Qwen3.8 27B = GPT-5.6 Luna compressed into 27B
How crazy is that? submitted by /u/kevinlch [link] [comments]
Reddit r/LocalLLaMA -
Limited sourcesModelsResearch
Artificial Analysis' Qwen3.8-27B benchmarks put it neck and neck with DeepSeek V4 and GPT-5.6 Luna Max
submitted by /u/anderspitman [link] [comments]
Reddit r/LocalLLaMA -
Limited sourcesModels
Qwen3.8 27B scores 52 on Artificial Analysis
Article URL: https://artificialanalysis.ai/models/qwen3-8-27b Comments URL: https://news.ycombinator.com/item?id=49334544 Points: 101 # Comments: 37
Hacker News Frontpage -
Limited sourcesModelsAgents & Dev
Cursor launches Origin, GitHub alternative
Aug 17, 2026 · Changelog Cursor can now host your code. Origin begins rolling out today in early beta on all paid plans. We're starting with the essentials, designed for agent scale: repos, pull requ…
Hacker News Frontpage -
Limited sourcesModels
Qwen3.8-27B at 256K on a 24GB RTX PRO 4000 SFF (432 GB/s): 50 tok/s with MTP
AI INFRASTRUCTURE I gave Qwen3.8's MTP drafter another 69.2 MiB of precision. Throughput fell from 50.44 to 37.02 tokens per second. That result sums up the whole experiment: the best local inference…
Hacker News Frontpage -
Limited sourcesModels
Qwen3.8-27B Q8_0 on Strix Halo is seriously impressive
Sorry for the slop, but I was impressed by this model as I have been testing Qwen3.8-27B Q8_0 locally on my ROG Flow Z13 (Ryzen AI Max+ 395, 128 GB unified memory) and this model was the only one who…
Reddit r/LocalLLaMA -
Limited sourcesModelsInfra
How many tokens/second output are you getting with Qwen3.8-27B?
Trying to get a feel for where I stand. If you can list your relevant hardware and model used, that would be awesome. Here's mine: Model: Qwen3.8-27B-heretic-ara, Q5_K_M GGUF T/s : ~30-32 t/sec (I th…
Reddit r/LocalLLaMA Aug 13 · Thu
-
Limited sourcesModels
llm-gemini 0.33
Release: llm-gemini 0.33 It's been a while since the last llm-gemini release. This version of the plugin adds support for today's Gemini 3.7 Flash release, plus gemini-3.6-flash , gemini-3.5-flash-li…
Simon Willison's Weblog -
ConfirmedModels
Previewing Ultrafast mode: GPT-5.6 Sol at up to 14X the speed
Lucas Ropek If you’ve ever found yourself wishing that ChatGPT was a little bit quicker on the uptake, OpenAI seems to be answering your prayers. The AI lab has rolled out a new mode called Ultrafast…
OpenAI News · TechCrunch AI -
ConfirmedModelsOpen Source
Introducing Gemini 3.7 Flash
Our most intelligent workhorse model yet for coding and agents. Tulsee Doshi Senior Director, Product Management, on behalf of the Gemini team Today, we’re building on the progress of our widely used…
Ars Technica AI · Google DeepMind Blog Aug 12 · Wed
-
ConfirmedModels
Introducing OlmoEarth embeddings: Custom embedding exports from OlmoEarth Studio for downstream analysis
📄 Tech Report: https://allenai.org/papers/olmoearth | 📊 Documentation: https://docs.olmoearth.allenai.org/embeddings | 💻 Learn more about OlmoEarth: https://allenai.org/olmoearth OlmoEarth Studio…
Hugging Face Blog Aug 11 · Tue
-
ConfirmedModels
Daybreak models are now available on AWS
OpenAI and AWS are making Daybreak cybersecurity capabilities available through Amazon Bedrock to support enterprise security workflows.
OpenAI News Aug 10 · Mon
-
Limited sourcesModels
Introducing Muse Glimmer
Introducing Muse Glimmer Meta are back in the open weights game! Muse Glimmer is a brand new 30B model under a clean Apache 2.0 license (a step up from the janky Llama licenses of old). They claim to…
Simon Willison's Weblog -
ConfirmedModels
Model ML completes finance work more efficiently with GPT-5.6 Sol
Model ML uses GPT-5.6 Sol to carry finance work from research and analysis through editable, traceable PowerPoint decks and Excel workbooks.
OpenAI News Aug 9 · Sun
-
Limited sourcesModels
Quoting Claude Opus 5 system prompt
Claude Fable 5 and Claude Mythos 5 were first released on June 9, 2026. On June 12, 2026, Anthropic suspended access to both models to comply with U.S. Department of Commerce export controls; the Dep…
Simon Willison's Weblog Aug 8 · Sat
-
Limited sourcesModelsAgents & Dev
Auto mode is now the default in Claude Code for Pro, Max, and Team plans
Auto mode is now the default in Claude Code for Pro, Max, and Team plans Anthropic are really confident in Claude Code's auto mode , to the point that they are making it the default setting for new s…
Simon Willison's Weblog Aug 7 · Fri
-
Limited sourcesModels
Moonlight & Mayhem (Raccoon Heist by Codex + GPT-5.6 Sol Ultra)
Moonlight & Mayhem (Raccoon Heist by Codex + GPT-5.6 Sol Ultra) On Wednesday I wrote about One-shotting a Raccoon Heist game using Claude Fable 5 , where I had Claude Fable 5 build a full working gam…
Simon Willison's Weblog Aug 6 · Thu
-
ConfirmedModels
Improving GPT‑5.6 Sol in ChatGPT—and expanding access to GPT-5.6 Luna for free users
ChatGPT introduces improved GPT-5.6 Sol with better accuracy and consistency, plus expanded access for free users and unlimited everyday chats with GPT-5.6 Luna.
OpenAI News Jul 30 · Thu
-
ConfirmedModels
Advancing the price-performance frontier with GPT-5.6
Explore lower GPT‑5.6 pricing for Luna and Terra—and how OpenAI’s more efficient models help enterprises deploy AI workflows at scale.
OpenAI News Jul 29 · Wed
-
ConfirmedModels
How GPT-5.6 fuses frontier intelligence with frontier efficiency
GPT-5.6 improves AI efficiency across models, inference, and agentic workflows, helping deliver more useful intelligence per dollar.
OpenAI News Jul 22 · Wed
-
ConfirmedModels
Introducing OpenAI Presence
Introducing OpenAI Presence, a proven enterprise AI agent platform that helps organizations deploy trusted voice and chat agents for customer and internal workflows.
OpenAI News Jul 21 · Tue
-
ConfirmedModels
Introducing the ChatGPT for small business program
OpenAI launches the ChatGPT for Small Businesses program, helping entrepreneurs build AI skills, automate work, and grow with ChatGPT Work.
OpenAI News -
ConfirmedModels
Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
Our newest Gemini models deliver the efficiency, latency, and reliability to build AI agents at scale. Tulsee Doshi Senior Director, Product Management, on behalf of the Gemini team Developers and cu…
Google DeepMind Blog