今日已更新 212 条资讯 | 累计 24871 条内容
关于我们

标签:#news

找到 9265 篇相关文章

AI 资讯

Ask HN: Are there good security benchmarks for LLMs?

I'm looking for this myself but figured it's good to have an actual discussion about this. I'm pretty new to the benchmarking side of LLMs. For example, I looked at eyeballvull [1]. It seems promising but I don't see wide support for example. I respect the author for still committing. A benchmark where an agent scans a repo in full is what I'm looking for. But then I also wondered: maybe there are others out there that I haven't been aware of yet. And, at the risk of potentially being flooded, f

2026-07-06 原文 →