今日已更新 230 条资讯 | 累计 24187 条内容
关于我们

标签:#hackernews

找到 8080 篇相关文章

AI 资讯

Ask HN: Are there good security benchmarks for LLMs?

I'm looking for this myself but figured it's good to have an actual discussion about this. I'm pretty new to the benchmarking side of LLMs. For example, I looked at eyeballvull [1]. It seems promising but I don't see wide support for example. I respect the author for still committing. A benchmark where an agent scans a repo in full is what I'm looking for. But then I also wondered: maybe there are others out there that I haven't been aware of yet. And, at the risk of potentially being flooded, f

2026-07-06 原文 →