Research 13 events
Category RSS ↗Today
-
Limited sourcesModelsAgents & Dev
Optimizing Qwen3.6 / Qwen3.8-27B on 16GB VRAM: Complete Benchmark Results and Setup Guide (~30-50tps at 32k to 72k context)
This post was made with AI. I tried to remove as much slop as possible and keep it straight to the point to save your time as I know how annoying AI slop posts can be, but I still wanted to retain al…
Reddit r/LocalLLaMA Yesterday
-
Limited sourcesResearch
we benchmark models nobody actually runs
qwen3.8-27b looks genuinely impressive on the benchmark tables - beating models many times its size on some of them. but those numbers come from bf16 weights, and nobody here is running a 27b at bf16…
Reddit r/LocalLLaMA -
Limited sourcesModelsResearch
Benchmarked Qwen3.8-27B on 4x RTX 3090
A while back I made a post about my 4x3090 rig in a Silverstone RV-02 . Check it out if you're a conoissuer of OG PC cases. With the incredible Qwen 3.8 27B release I ran benchmarks. So in case you a…
Reddit r/LocalLLaMA -
Limited sourcesOpen SourceAgents & Dev
Local agentic coding Benchmark : Qwen 3.8 27B (in many weights quants / cache quants / engine / reasoning effort) vs others.
In medium reasoning mode, it both scores higher than the 3.6 version, AND is very much more efficient (almost half requests needed, and a third less tokens generated) - at DeepSeek v4 Flash 3107 MXFP…
Reddit r/LocalLLaMA -
Limited sourcesModelsResearch
Artificial Analysis' Qwen3.8-27B benchmarks put it neck and neck with DeepSeek V4 and GPT-5.6 Luna Max
submitted by /u/anderspitman [link] [comments]
Reddit r/LocalLLaMA -
Limited sourcesResearch
[Paper] Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning
We introduce Mobius-v0, an architecture that comprises a globally shared Memory (FFN) that stores knowledge vectors and multiple Reasoners (Self-Attn) that iteratively achieve compositional reasoning…
Reddit r/LocalLLaMA -
Limited sourcesResearch
LLM's can't "jump" - a paper by Deepmind showing LLMs can't generate novel explanatory hypotheses
submitted by /u/juanviera23 [link] [comments]
Reddit r/LocalLLaMA Aug 13 · Thu
-
ConfirmedResearch
What We Learned by Reproducing 2,200 papers from ICML
Back in July, we ran a hackathon where more than 1,200 community members brought their own coding agents and tried to reproduce the papers published at ICML 2026, claim by claim. In 19 days, particip…
Hugging Face Blog Aug 11 · Tue
-
ConfirmedAgents & DevResearch
AMIE, our research medical AI system, demonstrates real-time clinical video consultation capabilities in a first-of-its-kind study.
Anil Palepu Research Lead When you visit a doctor, a consultation extends far beyond words — a physician notices a cough, observes gait, or registers visible signs of discomfort. Today, Google Resear…
Google AI Blog Aug 4 · Tue
-
ConfirmedResearchBusiness
Third-party cyber evaluations involving OpenAI models
OpenAI explains recent third-party cybersecurity evaluation incidents and outlines new safeguards to strengthen AI model testing and evaluation.
OpenAI News Jul 29 · Wed
-
ConfirmedResearch
How enabling two settings tripled our scores on the ARC-AGI-3 benchmark
How two API settings improved GPT-5.6 performance on ARC-AGI-3, boosting scores and efficiency by retaining reasoning and enabling compaction.
OpenAI News -
ConfirmedResearch
Accelerating scientific discovery with ChatGPT for Academic Researchers
OpenAI is giving 100,000 academic researchers free access to ChatGPT's most advanced AI models to accelerate scientific research, collaboration, and discovery.
OpenAI News Jul 21 · Tue
-
ConfirmedAgents & DevResearch
OpenAI and Hugging Face partner to address security incident during model evaluation
OpenAI and Hugging Face share early findings from a security incident during AI model evaluation, highlighting advanced cyber capabilities and lessons for defenders.
OpenAI News