本期重点:Vercel Sandbox新增Docker/Fuse支持,Cline集成Poolside的Laguna M.1免费模型;同时社区分享前端设计Skill测评和Fable自主判断调优技巧。

模型动态

美团LongCat 2.0模型权重正式发布

LongCat 2.0模型权重发布,提供INT8和FP8版本,对本地大模型社区是重要更新。Reddit讨论:https://www.reddit.com/r/LocalLLaMA/comments/1umo8zu/longcat_2_model_weights_have_been_published/[7]

产品工具

Vercel Sandbox新增Docker和Fuse支持,实现无约束MicroVM

Vercel的Sandbox现在支持docker和fuse,提供S3文件系统用于agent,对【AI agent基础设施】是重大更新。详情:https://x.com/rauchg/status/2073106884819841270[1]

Cline集成Poolside Laguna M.1免费模型

Cline新增免费模型Laguna M.1,225B参数,256k上下文,专为agentic coding优化。模型ID:poolside/laguna-m.1:free。来源:https://x.com/cline/status/2073086044665479378[2]

LlamaIndex发布集成Vercel Eve agent框架的模板

LlamaIndex推出模板集成Vercel Eveagent框架,提供只读文件系统工具和LiteParse解析功能,对Agent开发者和RAG场景有价值。https://x.com/llama_index/status/2073096981350605090[3]

技巧教程

@向阳乔木 实测5个前端设计Skill,给出详细评价

社区用户测评了5个流行前端设计Skill,认为ui-ux-pro-max一般,emil-design-eng动效最佳,web-design-guidelines规范强,taste-skillAI味最小,Anthropic自带frontend-design即将过时。[4]

Simon Willison分享Fable自主判断技巧:让模型自行调用子模型节省Token

Simon Willison建议让Fable模型自主判断何时调用更经济的子模型以节省token,并发布详细博客文章。项目地址:https://simonwillison.net/2026/Jul/3/judgement/[5][6]

硬件动态

用户实测Qwen 27B在4090+3090系统上的推理速度

Qwen 27B在4090+3090系统上实测达到解码50-90 tok/sprefill 1500-2200 tok/s,对本地部署选型有重要参考。Reddit:https://www.reddit.com/r/LocalLLaMA/comments/1umk3ax/qwen_27b/[8]

行业资讯

社区分析:RAG场景中prefill速度比decode更重要

Reddit深入分析指出,在RAG场景中【prefill速度】比decode更关键,并讨论Strix Halo等硬件的瓶颈。对构建RAG系统有洞察价值。https://www.reddit.com/r/LocalLLaMA/comments/1umlqwn/for_rag_specifically_prefill_speed_matters_more/[9]