MiniMax MiniMax H3 Max探索快于实时的视频生成体验 MiniMax称,fal通过后训练与推理优化,让H3 Max实现【快于实时的视频生成】,有望支持持续直播与互动世界。https://x.com/MiniMax_AI/status/2093778769903534294[1][2]
DeepSeek / Zhipu 社区实测DeepSeek V4 Flash与GLM 5.3 Flash代码能力 社区在双DGX Spark环境下比较DeepSeek V4 Flash 0731与GLM 5.3 Flash的HumanEval表现,为本地模型选型提供参考。https://www.reddit.com/r/LocalLLaMA/comments/1w215qm/humaneval_benchmark_for_deepseek_v4_flash_0731_vs/[4]
Nemotron-3.5-Lightning被压缩至11.77 GiB适配16GB显存 社区通过调整量化时的tensor宽度,并在推理阶段恢复原尺寸,将模型压至11.77 GiB,解决16GB设备适配问题。https://www.reddit.com/r/LocalLLaMA/comments/1w21d86/nemotron35lightning_at_1177_gib_a_16_gb_option/[5]
Qwen 本地用户讨论Qwen3.8 Next与27B量化版本取舍 讨论显示,Qwen3.8 Next UD IQ1_S速度明显更慢,准确率约七成以上;用户需权衡量化等级、【推理速度】与精度。https://www.reddit.com/r/LocalLLaMA/comments/1w21d1p/any_reason_to_use_qwen_38_next_ud_iq1_s_over_qwen/[7]
Qwen 社区公布Tenstorrent QuietBox 2运行Qwen3.7-27B基准 新帖整理QuietBox 2及P300C配置运行Qwen3.7-27B的数据,补充Tenstorrent本地推理硬件的公开性能参考。https://www.reddit.com/r/LocalLLaMA/comments/1w21d86/tenstorrent_qwen372b_benchmarks/[6]