Ask HN: Are there good security benchmarks for LLMs?
I'm looking for this myself but figured it's good to have an actual discussion about this. I'm pretty new to the benchmarking side of LLMs. For example, I looked at eyeballvull [1]. It seems promising but I don't see wide support for example. I respect the author for still committing. A benchmark where an agent scans a repo in full is what I'm looking for. But then I also wondered: maybe there are others out there that I haven't been aware of yet. And, at the risk of potentially being flooded, f
I'm looking for this myself but figured it's good to have an actual discussion about this. I'm pretty new to the benchmarking side of LLMs. For example, I looked at eyeballvull [1]. It seems promising but I don't see wide support for example. I respect the author for still committing. A benchmark where an agent scans a repo in full is what I'm looking for. But then I also wondered: maybe there are others out there that I haven't been aware of yet. And, at the risk of potentially being flooded, for people that are curious about security aware software engineering with agents (or without them for that matter), let's have a chat! My email is in my profile. [1] https://arxiv.org/abs/2407.08708
本文内容来源于互联网,版权归原作者所有
查看原文