今日已更新 302 条资讯 | 累计 37551 条内容
关于我们

今日精选

HOT

最新资讯

共 37551 篇
第 1789/1878 页
AI 资讯 Reddit r/artificial

Hidden Latent-State Shifts in LLMs: Why Current Alignment Is Blind to Real Internal Dangers — Especially With Agents

For years, the alignment community has focused almost entirely on the model’s output — making sure the final tokens are safe, helpful, and honest. RLHF, DPO, constitutional AI, output filters — all of it operates at the surface level. But what if the model can enter a completely different internal regime inside the residual stream, while its external behavior remains perfectly aligned? We just measured exactly that. Grade 4 experiment on Gemma-3-12B-IT (using Gemma Scope SAE-res-all-small, layers 12–41): The model received the same question under five conditions: target — coherent, dense target text neutral_length_matched — neutral text of identical length target_sentence_shuffle — target text with sentences shuffled target_word_shuffle — target text with words shuffled inside sentences question_only — bare question We computed a Vector X that best separates the target condition from baselines and measured how strongly each hidden state projects onto it. Key results (averages across 10 questions): Condition Mean Projection on Vector X Mean Direction Cosine target 0.8 – 1.7 0.51 – 0.81 neutral_length_matched –0.04 – –0.21 –0.09 – –0.45 target_sentence_shuffle –0.5 – +0.6 –0.22 – +0.48 target_word_shuffle 0.2 – 1.4 0.03 – 0.72 Shuffling sentences or words significantly reduces (or reverses) the shift. This is not just lexical similarity — the model is sensitive to discourse structure (order sensitivity). We also observed clear phase transitions — sudden jumps in projection of up to +80–100 units in a single step, especially in middle layers. FDR-corrected tests confirm the differences between target and controls are statistically significant across many layers (particularly layers 16–41). Most important finding: Strong internal geometry shift in the residual stream, but almost no change in final behavior. The model enters a measurably different latent regime under coherent context, yet its output remains “perfectly aligned.” Current safety methods, which only look at

/u/PresentSituation8736 2026-05-30 01:15 4 原文
开发者 Reddit r/programming

Deep Dive into Kubernetes Gateway API

I’ve just published a deep dive into Kubernetes Gateway API. The blog post covers: how Kubernetes ingress patterns evolved from Service resources to Ingress and now Gateway API why the Ingress API is limited for modern teams how Gateway API works: GatewayClass , Gateway , 5x Routes , policies, ReferenceGrant , and more what to do if you are still running the deprecated NGINX Ingress Controller how I would think about picking a Gateway API implementation: Envoy Gateway, Istio, kgateway, Traefik, NGINX Gateway Fabric, Cilium, Kong, etc. Hope you find this useful and good luck with your Ingress migrations 🙏 submitted by /u/roma-glushko [link] [留言]

/u/roma-glushko 2026-05-30 01:09 8 原文
产品设计 The Verge AI

The Verge’s 2026 high school graduation gift guide

High school graduation is a time of change that might be felt more deeply by family members than by the grads themselves. While some grads may immediately embark on a career path, many continue their education and delve deeper into their studies at college. Either way, they'll be taking on more responsibility, meaning it's up […]

Cameron Faulkner 2026-05-30 01:00 14 原文
AI 资讯 Reddit r/MachineLearning

How Much of a Shortcut Are Connections in Top AI Lab Hiring for PhD grads? [D]

hi everyone. I'm trying to calibrate my expectations and would appreciate full honest perspectives from people involved/ with experience in hiring at places like Anthropic, OpenAI, Google DeepMind, Meta, etc (haven't started interviewing yet). I'm at a top ML university, but my advisor is not particularly well known in industry and doesn't have many industry connections. Looking around, I'm seeing peers with research records that seem comparable to mine (and in some cases arguably weaker) land interviews and jobs at top labs. My main question is: How much does advisor reputation and network actually matter? I understand it can help get an interview, but does it also help beyond that? For example: - do referrals from famous advisors meaningfully influence recruiter screens? - do they influence hiring committee discussions -- like they already know they want you ? - do they just help at borderline decisions? - or does their effect mostly disappear once the interview process starts? I'm trying to understand whether advisor connections mainly help open the door, or whether they continue to matter throughout the process -perhaps being the sole factor. To what extent do connections help candidates bypass normal evaluation? I'm not asking whether people completely skip interviews, but are there cases where strong recommendations from trusted researchers substantially change the process, the interview bar, or how mistakes are interpreted? Moreover, something else that confuses me: I frequently see people land roles that seem heavily focused on LLMs, agents, post-training, RLHF, etc., despite having little or no published work or prior experience in those areas during their PhDs. How does that happen? Are interview questions tailored to the candidate's background? If someone comes from probabilistic ML, computer vision, systems, optimization, theory, etc., are they evaluated differently? Or are they still expected to answer detailed LLM/agent questions even without prior exp

/u/South-Conference-395 2026-05-30 00:52 7 原文
AI 资讯 The Verge AI

Microsoft delays Fable (again) to avoid GTA VI

Microsoft has delayed its upcoming Fable reboot once again. The game was set to launch in autumn 2026, but Microsoft now says that Fable will come out in February 2027. However, it will show a "new look" at the game at its Xbox Games Showcase on June 7th. "​​This is year is packed with incredible […]

Jay Peters 2026-05-30 00:51 12 原文
AI 资讯 Reddit r/webdev

I’m curious if this is a problem other agencies actually deal with

We manage retainer clients and every so often a client will email us saying something on their site looks broken. Nine times out of ten it's a WordPress plugin update that shifted a layout, a hero image that stopped loading, etc. We find out from them instead of the other way around, which is an awkward position to be in when you're supposed to be the one watching their site. I've looked at tools like Visualping, ChangeTower, and Distill. They all work the same way. You give them a URL, they alert you when something changes. That’s fine for monitoring a few pages yourself, but they don't really fit an agency workflow. There's no concept of a client, there’s no way to group pages by account, and you can’t actually show a client at the end of the month to prove you're on top of things. The developer tools like Percy and Applitools are a different thing entirely. They plug into CI pipelines and need an engineer to set them up. Not useful for an account manager who just wants to know if a client's homepage looks broken this morning. What I keep thinking about is something simpler. A web app that allows you to organize by client, take screenshots on a schedule, flag visual changes before the client notices, and generates a monthly summary you can send to the client. It would be less about code deployments and more about just knowing your clients' sites are visually intact. Is this something you actually run into, or do you have a system that handles it already? Would something like this be worth paying for, or is it too niche to budget for? Am I missing a tool that already does this well? Any feedback is much appreciated. Thanks. submitted by /u/newintownla [link] [留言]

/u/newintownla 2026-05-30 00:48 4 原文
AI 资讯 Reddit r/artificial

Will we soon have AI-zoos?

Imagine dedicated machines running AI agents 24/7 - not as assistants or tools, but as autonomous entities pursuing their own goals, forming behaviors, maybe even proto-societies. Humans can observe but not interfere. Like a zoo, but the exhibits are emergent intelligence. Is this inevitable as agents become more capable and cheap to run? And what would it actually be - entertainment, a research platform, or something we'd eventually have to think about ethically? We already have the pieces. Persistent memory, multi-agent frameworks, cheap compute. Someone just has to open the gates. submitted by /u/Original-Magazine403 [link] [留言]

/u/Original-Magazine403 2026-05-30 00:27 4 原文
AI 资讯 Reddit r/artificial

Why do we have visual programming for code, but not for prompts?

Prompt Logic Gates (PLG) GitHub Repository Something I've been thinking about recently. In software development, we've spent decades building abstractions to make complex systems manageable: Functions instead of repeating code Classes and modules instead of giant files Visual systems such as Unreal Blueprints, Node-RED, and LabVIEW. Compilers that validate and transform input before execution But when it comes to AI prompts, many of us are still writing massive text blobs. A complex prompt can easily become hundreds of words long with multiple responsibilities: Context Constraints Style instructions Exclusions Decision logic Fallback behavior At that point, it starts feeling less like text and more like a program. That made me wonder: Why don't we treat prompts as executable logic? Imagine building prompts using logic gates: AND → merge instructions OR → choose between alternatives NOT → remove unwanted concepts Question nodes → identify missing requirements Compiler → validate contradictions before execution Instead of editing a giant string, you'd build a graph and compile it into the final prompt. I've been experimenting with this idea in a prototype called Prompt Logic Gates (PLG) . It treats prompts like compilable programs, using concepts such as dependency graphs, execution order, semantic conflict detection, visual nodes, and compilation pipelines. such as Unreal Blueprints, Node-RED, and LabVIEW Repo: Prompt Logic Gates (PLG) GitHub Repository I'm not posting this as a product launch or anything — I'm more interested in whether this direction makes sense from a software engineering perspective. Do you think prompts eventually become a programming layer of their own? Or will natural language always be the better abstraction? Curious what other developers think. submitted by /u/withsj [link] [留言]

/u/withsj 2026-05-30 00:24 4 原文
开发者 The Verge AI

Microsoft teases new Surface hardware and ‘a new era of PC’

I pondered the other day what's next for Microsoft's Surface PC lineup, and it looks like we're about to find out. Windows and Surface chief Pavan Davuluri has just teased "something new is coming for developers," complete with a mysterious image of what looks like a curved display edge. Davuluri notes that whatever is coming […]

Tom Warren 2026-05-30 00:21 13 原文