<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Benchmarks and evaluation — USASI news</title>
    <link>https://unitedstatesofamericasuperintelligence.com/hubs/evaluation/</link>
    <atom:link href="https://unitedstatesofamericasuperintelligence.com/hubs/evaluation/feed.xml" rel="self" type="application/rss+xml"/>
    <description>Dated, sourced news items about the organizations and records featured in the USASI Benchmarks and evaluation hub. Independent project. Not a United States government website.</description>
    <language>en-us</language>
    <lastBuildDate>Thu, 01 Oct 2026 12:00:00 GMT</lastBuildDate>
    <item>
      <title>MLCommons publishes MLPerf Inference v6.1 results with two new tests</title>
      <link>https://unitedstatesofamericasuperintelligence.com/news/mlcommons-publishes-mlperf-inference-v6-1/</link>
      <guid isPermaLink="true">https://unitedstatesofamericasuperintelligence.com/news/mlcommons-publishes-mlperf-inference-v6-1/</guid>
      <pubDate>Thu, 01 Oct 2026 12:00:00 GMT</pubDate>
      <category>release</category>
      <description>MLCommons published results for version 6.1 of the MLPerf Inference benchmark suite on September 16, 2026. The round added two tests: an end-to-end retrieval-augmented generation test that runs a question-answering pipeline of several separate models, and an Edge Agentic Inference test built around multi-turn workloads such as agentic coding. Speculative decoding is now allowed in the interactive scenario for two of the inference benchmarks. Results are published on MLCommons' datacenter and edge results pages.</description>
    </item>
    <item>
      <title>Terminal-Bench 4.0 released as a new major version of the continuous benchmark</title>
      <link>https://unitedstatesofamericasuperintelligence.com/news/terminal-bench-4-0-released/</link>
      <guid isPermaLink="true">https://unitedstatesofamericasuperintelligence.com/news/terminal-bench-4-0-released/</guid>
      <pubDate>Tue, 29 Sep 2026 12:00:00 GMT</pubDate>
      <category>release</category>
      <description>The Terminal-Bench team published version 4.0 of its command-line agent benchmark on August 28, 2026, a month after announcing that it would be maintained as a continuous benchmark with semantic versioning. The release removed eight tasks, fixed nineteen, and set a flat eight-hour agent timeout for every task; it is a major version because these changes require re-running trials. It runs through the Harbor framework.</description>
    </item>
  </channel>
</rss>
