MEPPP 热点收录
@meppp_hot
收录
我们正在发布AutoResearchExam ,这是开放式机器学习和工程任务的基准。我们的基准涵盖七个研究领域,包括模型培训、数据管理、人工智能安全性和可解释性。https://benchmarks.bespokelabs.ai/autoresearchexam/
在每项任务中,我们给代理24小时使用CPU或GPU机器来开发和改进他们的解决方案通过实验和反馈。我们通过综合得分来衡量速度和质量。我们的基准有一个独特的功能:测试客服代表是否可以改进他们从未见过的数据。我们发现,人工智能研究代理商在尝试改进时往往过度拟合。我们在前沿看到了一个有趣的直接比较: Astra从最强者开始,并保持领先优势,19小时,但Fable 5.1在最后几个小时赶上并获得最佳表现。Qwen3.8 MAX、Gemini 3.8 Flash和Grok 4.6都位于高性价比的帕累托前沿,提供了强大的选择以更低的API预算。
Anthropic的Opus和Fable几乎保留了所有验证隐藏测试的性能,差距为1.1 %和2.9 %。Astra相对于Sol的改进也延伸到了泛化,该差距从6.9%下降到1.7%。( 1/n )
在每项任务中,我们给代理24小时使用CPU或GPU机器来开发和改进他们的解决方案通过实验和反馈。我们通过综合得分来衡量速度和质量。我们的基准有一个独特的功能:测试客服代表是否可以改进他们从未见过的数据。我们发现,人工智能研究代理商在尝试改进时往往过度拟合。我们在前沿看到了一个有趣的直接比较: Astra从最强者开始,并保持领先优势,19小时,但Fable 5.1在最后几个小时赶上并获得最佳表现。Qwen3.8 MAX、Gemini 3.8 Flash和Grok 4.6都位于高性价比的帕累托前沿,提供了强大的选择以更低的API预算。
Anthropic的Opus和Fable几乎保留了所有验证隐藏测试的性能,差距为1.1 %和2.9 %。Astra相对于Sol的改进也延伸到了泛化,该差距从6.9%下降到1.7%。( 1/n )
查看原文
原帖正文
We are releasing AutoResearchExam, a benchmark on open-ended machine learning and engineering tasks. Our benchmark covers seven research areas including model training, data curation, AI safety and interpretability.https://benchmarks.bespokelabs.ai/autoresearchexam/In each task, we give agents 24 hours with a CPU or GPU machine to develop and improve their solutions through experiments and feedback. We measure both speed and quality with a combined score. Our benchmark has a unique feature: testing if agents create improvements that hold up on data they never see. We find that AI research agents often overfit as they try to improve.We see an interesting head-to-head comparison at the frontier: Astra starts the strongest and holds the lead for up to 19 hours but Fable 5.1 catches up and gets the top performing spot in the final hours.Qwen3.8 Max, Gemini 3.8 Flash and Grok 4.6 all sit on the cost-performance Pareto frontier, giving strong options at lower API budgets.Anthropic's Opus and Fable retain nearly all their validation performance on hidden tests, with gaps of 1.1% and 2.9%. Astra's improvement over Sol extends to generalization too, with that gap falling from 6.9% to 1.7%. (1/n)
引用自 X
原帖地址:https://x.com/i/status/2097757256783970713
原帖:
本站媒体副本
0 条评论
登录后可以参与讨论。
成员登录还没有评论。第一条认真回应会很重要。