研究 · 评测 13 个事件
本分类 RSS ↗今天
-
有限来源模型发布Agent · 编程
Optimizing Qwen3.6 / Qwen3.8-27B on 16GB VRAM: Complete Benchmark Results and Setup Guide (~30-50tps at 32k to 72k context)
This post was made with AI. I tried to remove as much slop as possible and keep it straight to the point to save your time as I know how annoying AI slop posts can be, but I still wanted to retain al…
Reddit r/LocalLLaMA 昨天
-
有限来源研究 · 评测
我们在评测根本没人真正跑的模型译
qwen3.8-27b looks genuinely impressive on the benchmark tables - beating models many times its size on some of them. but those numbers come from bf16 weights, and nobody here is running a 27b at bf16…
Reddit r/LocalLLaMA -
有限来源模型发布研究 · 评测
在 4 张 RTX 3090 上实测 Qwen3.8-27B译
A while back I made a post about my 4x3090 rig in a Silverstone RV-02 . Check it out if you're a conoissuer of OG PC cases. With the incredible Qwen 3.8 27B release I ran benchmarks. So in case you a…
Reddit r/LocalLLaMA -
有限来源开源Agent · 编程
本地 Agentic 编程评测:Qwen 3.8 27B(多种权重量化/缓存量化/引擎/推理档位)对比其他模型译
In medium reasoning mode, it both scores higher than the 3.6 version, AND is very much more efficient (almost half requests needed, and a third less tokens generated) - at DeepSeek v4 Flash 3107 MXFP…
Reddit r/LocalLLaMA -
有限来源模型发布研究 · 评测
Artificial Analysis 评测:Qwen3.8-27B 与 DeepSeek V4、GPT-5.6 Luna Max 并驾齐驱译
submitted by /u/anderspitman [link] [comments]
Reddit r/LocalLLaMA -
有限来源研究 · 评测
[论文] Intern-S2-Mobius:知识与推理解耦的基础模型译
We introduce Mobius-v0, an architecture that comprises a globally shared Memory (FFN) that stores knowledge vectors and multiple Reasoners (Self-Attn) that iteratively achieve compositional reasoning…
Reddit r/LocalLLaMA -
有限来源研究 · 评测
LLM 不会「跳跃」——DeepMind 论文指出 LLM 无法产生新颖的解释性假说译
submitted by /u/juanviera23 [link] [comments]
Reddit r/LocalLLaMA 8月13日 · 周四
-
已确认研究 · 评测
复现 2200 篇 ICML 论文教会我们的事译
Back in July, we ran a hackathon where more than 1,200 community members brought their own coding agents and tried to reproduce the papers published at ICML 2026, claim by claim. In 19 days, particip…
Hugging Face Blog 8月11日 · 周二
-
已确认Agent · 编程研究 · 评测
医疗研究 AI 系统 AMIE 在首创研究中展示实时临床视频问诊能力译
Anil Palepu Research Lead When you visit a doctor, a consultation extends far beyond words — a physician notices a cough, observes gait, or registers visible signs of discomfort. Today, Google Resear…
Google AI Blog 8月4日 · 周二
-
已确认研究 · 评测商业动态
涉及 OpenAI 模型的第三方网络安全评估译
OpenAI explains recent third-party cybersecurity evaluation incidents and outlines new safeguards to strengthen AI model testing and evaluation.
OpenAI News 7月29日 · 周三
-
已确认研究 · 评测
开启两个设置让我们的 ARC-AGI-3 得分翻三倍译
How two API settings improved GPT-5.6 performance on ARC-AGI-3, boosting scores and efficiency by retaining reasoning and enabling compaction.
OpenAI News -
已确认研究 · 评测
用 ChatGPT 学术研究者版加速科学发现译
OpenAI is giving 100,000 academic researchers free access to ChatGPT's most advanced AI models to accelerate scientific research, collaboration, and discovery.
OpenAI News 7月21日 · 周二
-
已确认Agent · 编程研究 · 评测
OpenAI 与 Hugging Face 联手处置模型评估期间的安全事件译
OpenAI and Hugging Face share early findings from a security incident during AI model evaluation, highlighting advanced cyber capabilities and lessons for defenders.
OpenAI News