<rss xmlns:atom="http://www.w3.org/2005/Atom" version="2.0">
  <channel>
    <title>arxiv - 标签 - 搬砖程序员带你飞</title>
    <link>https://lpflpf.cn/tags/arxiv/</link>
    <description>arxiv - 标签 - 搬砖程序员带你飞</description>
    <generator>Hugo -- gohugo.io</generator><language>zh-CN</language><managingEditor>lipengfei634626165@163.com (lpflpf)</managingEditor>
      <webMaster>lipengfei634626165@163.com (lpflpf)</webMaster><lastBuildDate>Sat, 15 Aug 2026 15:00:00 &#43;0800</lastBuildDate><atom:link href="https://lpflpf.cn/tags/arxiv/" rel="self" type="application/rss+xml" /><item>
  <title>评测 Agent 别只看分数：美团这篇论文把 7 个模型、36 个任务拆开看，发现它们更像「工程优化器」而不是「研究者」</title>
  <link>https://lpflpf.cn/posts/paper-beyond-final-scores/</link>
  <pubDate>Sat, 15 Aug 2026 15:00:00 &#43;0800</pubDate>
  <author>lpflpf</author>
  <guid>https://lpflpf.cn/posts/paper-beyond-final-scores/</guid>
  <description><![CDATA[<div class="featured-image">
        <img src="https://cdn.lpflpf.cn/covers/paper-beyond-final-scores.png" referrerpolicy="no-referrer">
      </div>看 Agent 能力，你习惯看什么？benchmark 最终分数，对吧。 但一份最终分数回答不了这几个问题：这个 Agent 的进步发生在哪个环节？它积累的经验有没有真]]></description>
</item>
<item>
  <title>从 DeepSeek Harness 到 arXiv：一周冒出 5 篇 Harness 论文，Agent 执行层正在成为研究主战场</title>
  <link>https://lpflpf.cn/posts/paper-review-agent-harness/</link>
  <pubDate>Sat, 15 Aug 2026 14:00:00 &#43;0800</pubDate>
  <author>lpflpf</author>
  <guid>https://lpflpf.cn/posts/paper-review-agent-harness/</guid>
  <description><![CDATA[<div class="featured-image">
        <img src="https://cdn.lpflpf.cn/covers/paper-review-agent-harness.png" referrerpolicy="no-referrer">
      </div>上周 DeepSeek Harness 开源，2.3 万 star 一天到账。当时我说&quot;模型是商品，生态才是护城河&quot;——结果这周 arXiv 直接用论文回应了：Harness 不]]></description>
</item>
</channel>
</rss>
