今日已更新 255 条资讯 | 累计 37160 条内容
关于我们

今日精选

HOT

最新资讯

共 37160 篇
第 1738/1858 页
AI 资讯 Reddit r/MachineLearning

I built mlx-Chronos — a community benchmark leaderboard for local LLM engines on Apple Silicon (oMLX, Rapid-MLX, mlx-lm, Ollama) [P]

Hey! I'm a CS student and I got tired of not being able to compare MLX inference engines properly — every benchmark out there is either made by the engine's own developers, runs on an M3 Ultra nobody has, or just shows tok/s with zero context. So I built mlx-Chronos — a small open source CLI tool that runs a standardized benchmark protocol on your Mac and lets you submit your results to a shared community leaderboard. What it measures: Cold and cached TTFT (Time to First Token), with a proper methodology — unique prompts per trial, cache priming, no interleaved phases Throughput (tok/s), with mean/stddev/min/max across repeated trials Engine process RSS and system RAM peak, sampled continuously during inference Thermal state and hardware info Supported engines: oMLX, Rapid-MLX, mlx-lm, Ollama (MLX backend) The leaderboard is basically empty right now since I only have an M2 8GB. Would love results from M3 Max, M4, M4 Ultra, or anything with more RAM — that's where things get actually interesting. → Leaderboard: https://igurss.github.io/mlx-chronos → GitHub: https://github.com/igurss/mlx-chronos → Install: pip install mlx-chronos It's early, the methodology is documented (there's a methodology.md if you want to pick it apart), and I'm 100% open to feedback, contributions, and getting told what I'm doing wrong. The goal is just to have one place where you can compare engines on your specific hardware instead of trusting someone else's numbers. submitted by /u/igor__004 [link] [留言]

/u/igor__004 2026-05-31 16:26 6 原文
AI 资讯 Reddit r/webdev

Have DIY website builders had an effect on the web-dev industry?

I've recently been getting reacquainted with frontend dev after having spent the last few years focusing mostly on design. I plan to do a combination of both again from now on. The last full-time job I was at we used drag-and-drop website builders that seem to be all the rage these days. I gotta say it was a truly painful experience. The websites we would handover would be a complete mess, clients would have to edit columns, widgets, layouts themselves which just felt like a recipe for disaster. We would often have to duplicate various widgets such as image sliders, one of mobile and one for desktop. I'd then have to explain to clients that they needed to update them twice, with the same content. Also, a lot of times I'd have to set up HTML widgets because of how limited the software was, so I'd also have to explain to clients that they'd sometimes need to edit actual code when simply updating content on their site. But the concerning thing is, I've had a number of other job interviews at marketing agencies lately and it seems like EVERYONE is doing the same thing. I also sometimes hear on podcasts people describing default WordPress as being 'old fashioned', as though the future of web dev is just doing everything with these hideous website builders WTF 🤮 Are these people mainly just taking the basic / shitty web dev projects while leaving the proper ones to those who know how to do it properly? or have a lot of developers lost out on work because of these? submitted by /u/Weekly_Frosting_5868 [link] [留言]

/u/Weekly_Frosting_5868 2026-05-31 16:07 5 原文
AI 资讯 Smashing Magazine

June Is For Exploring (2026 Wallpapers Edition)

Let’s kick off June — and the beginning of summer — with some fresh inspiration! Artists and designers from across the globe once again tickled their creativity to welcome the new month with a new collection of desktop wallpapers. Enjoy!

hello@smashingmagazine.com (Cosima Mielke) 2026-05-31 16:00 4 原文
AI 资讯 Reddit r/artificial

Anyone tried using AI models to screen candidates?

I used these two prompts on all AI apps to figure out who to vote for in the CA primaries: If you were running for governor of California, what will your big policies be ⁠Out of the candidates that are running in June election, who aligns closest to those policies Gemini, claude, chatgpt all ranked Matt Mahan (Democrat) as #1 Grok chose Steve Hilton (Republican) thoughts on AI use for voting decisions? submitted by /u/No_Mall_7299 [link] [留言]

/u/No_Mall_7299 2026-05-31 15:04 4 原文
AI 资讯 Dev.to

103. Agent Memory: Short-Term, Long-Term, and Episodic

Agent Memory: Short-Term, Long-Term, and Episodic Main Thumbnail Image Prompt: A human brain cross-section illustration in neon tones on dark background. Three regions clearly demarcated and labeled. The hippocampus region glows blue, labeled "Episodic Memory: what happened." The prefrontal cortex glows orange, labeled "Working Memory: what I'm doing now." A network of distributed nodes glows green, labeled "Semantic Memory: what I know." Arrows show information flowing between regions. Scientific but accessible, the memory architecture made neural and visual. Memory Architecture Diagram Image Prompt: Four storage boxes arranged vertically on dark background. Top: "In-Context Window (Working Memory)" — fastest, smallest, temporary, shown as RAM chip icon. Second: "External Vector Store (Semantic Memory)" — fast retrieval, persistent, shown as cylinder with search icon. Third: "Key-Value Store (Episodic Memory)" — structured facts, shown as database icon. Bottom: "Fine-Tuned Weights (Procedural Memory)" — slowest to update, most permanent, shown as brain with lock. Arrows showing read/write speeds between boxes. Clean, technical, the hierarchy is the insight. Memory Retrieval Flow Image Prompt: A query arrives at an agent on the left. Four parallel arrows go right to four memory sources: conversation history (short chat bubbles), vector database (semantic search visualization), structured database (table icon), model weights (brain icon). Each source returns relevant items. A "Memory Fusion" box on the right combines the results. The agent sees an enriched context. The retrieval from multiple stores is the architecture. Every conversation with an LLM starts from zero. You explain your project. You explain your preferences. You explain your constraints. You spend five minutes providing context. You come back tomorrow. You do it all again. The model remembers nothing between sessions. The context window closes. The state is gone. Every interaction is the agent's first

Akhilesh 2026-05-31 14:53 7 原文
AI 资讯 Reddit r/artificial

Robot foundation models keep hiding behind fine-tuning numbers. Wall-OSS-0.5 is trying a different approach

Most robot foundation model demos are hard to interpret because the impressive number usually comes after task-specific fine tuning. Wall-OSS-0.5, a new open-source VLA release from X Square Robot, is interesting because the report tries to measure what the pretrained checkpoint can do before that extra adaptation step. The setup is a 4B vision-language-action model built around a 3B VLM backbone plus action-generation components. According to the report, the pretrained checkpoint was evaluated on a 17-task real-robot suite without task-specific fine tuning. Four tasks crossed 80 task progress: block sorting, fruit sorting, ring stacking, and a held-out deformable task, rope tightening. The part that seems more important than the raw score is the framing. In language models, nobody would accept only a fine-tuned downstream score as evidence that pretraining worked. With robots, that has been much harder because the evaluation is physical, slow, embodiment-dependent, and expensive. A real-robot zero-shot suite is a useful step toward asking the same question directly: does pretraining itself produce executable behavior, or is it mostly a better initialization? The method is also trying to solve a specific training problem. Continuous action losses are useful for execution, but the paper argues they do not send a strong enough learning signal into the VLM backbone by themselves. Their recipe combines action-token cross entropy, multimodal cross entropy, and flow matching in one stage, using the discrete action-token path as a gradient bridge into the backbone while flow matching handles continuous actions at deployment time. For reference, the code is at https://github.com/X-Square-Robot/wall-x , the paper is at https://x2robot.com/api/files/file/wall_oss_05.pdf , the project page is https://x2robot.com/oss#resources , and the Hugging Face org is https://huggingface.co/x-square-robot . The caveat is obvious but important. Zero-shot still does not solve the hardest man

/u/breadislifeee 2026-05-31 14:50 7 原文
开发者 Dev.to

The Most Used Technology in the World Has Zero Marketing and Product People

174 million smart TVs, most of which run Linux. 3.9 billion Android phones. Zero marketing. Tonight, somewhere around the world, a person will press the power button on their Samsung TV. A proprietary Samsung logo will appear. A polished menu will load. They will open Netflix, scroll through recommendations, and pick a movie. They will never know that every frame they see is being scheduled, managed, and rendered by a Linux kernel, the invisible engine that sits between apps and hardware. They will then reach for their Android phone to check something on social media. Another Linux kernel. If they are sitting in a Tesla, the touchscreen showing their charging status is running yet another Linux kernel. The “year of the Linux desktop” debate has been running for two decades. Entire forums exist to argue about whether 2025, 2026, or 2027 will finally be the year Linux takes over the PC market.

Himanshu Kumar 2026-05-31 14:40 12 原文
AI 资讯 Dev.to

The Principle of Least Privilege: Operational Speed's Security Cost

The Principle of Least Privilege: Operational Speed's Security Cost While developing a production ERP, delayed shipment reports were always a headache. One of the main reasons behind incomplete reports was the complexity of privilege layers in the system and, often, excessive permissions granted. In this post, I will delve into the costs we pay when we stretch security boundaries in an effort to gain operational speed. The principle of least privilege is more than just a security concept; it's critically important for operational efficiency and system stability. In this article, I will explain the impact of the principle of least privilege on operational speed, the security risks it entails, and how I've tried to strike this balance with concrete examples from my practical experience. My goal is to move beyond superficial definitions and dive deep into this topic based on my real-world field experiences, providing actionable insights to readers. Why Does the Principle of Least Privilege Seem to Hinder Operational Speed? The general tendency is to provide instant access to all relevant tools and data to speed up a task. This can be appealing, especially in an emergency or before a critical delivery. However, the Principle of Least Privilege (PoLP) advocates the opposite: a user or system component should have the absolute minimum privileges required to perform its task. This might initially seem to slow down operational processes. For example, a development team having unlimited SELECT rights to a production database might facilitate running an urgent query. However, the same developer could accidentally run UPDATE or DELETE commands, causing serious damage to the system. Such an incident, instead of speeding up a query in the short term, could lead to hours of downtime and data loss. This is where the long-term risk posed by operational speed, which PoLP is thought to hinder, becomes apparent. Another example is a system administrator frequently using the sudo su co

Mustafa ERBAY 2026-05-31 14:35 11 原文
AI 资讯 Dev.to

Your AI Sucks at Math. Fix It With One Command.

You've seen this before. You ask your AI agent: "Find ∫ x·e^x dx" It confidently replies: e^x + C , complete with a plausible-looking derivation. You nod. Then you check — the correct answer is (x−1)·e^x + C . It was wrong by a mile, and you almost shipped it. This is the fundamental problem with AI math today: LLMs can talk, but they can't verify their own work. They sound convincing while being catastrophically wrong. And the more complex the problem, the better the hallucination. Math.skill changes that. It's an open-source mathematical reasoning skill for AI agents — install it, and your agent stops guessing and starts verifying. What Makes It Different Typical AI Math Plugin Math.skill Workflow Prompt → LLM → answer Prompt → 7-step pipeline → ≥2 verifications → answer Verification None Answer blocked if verification fails Open problems Might hallucinate a "solution" Honestly says "this is unsolved" Error recovery No mechanism Auto-backtrack, fix, recompute, re-verify The core differentiator: a verification engine that runs at least 2 of 11 independent checks on every answer. No answer leaves the pipeline unverified. Period. The 7-Step Pipeline Every problem flows through this: Step What Happens Why It Matters 1. Parse Extract conditions, goals, variables, implicit domain constraints Catches misread problems before they waste your time 2. Model Build formal representation: equation, function, matrix, probability space, etc. Prevents building the wrong mathematical structure 3. Select Choose the optimal method from 30+ strategies Avoids brute-forcing when elegance exists 4. Solve Step-by-step with mathematical justification at every transformation Full traceability — nothing hidden 5. Verify Apply ≥2 of 11 independent verification methods The differentiator — catches what LLMs miss 6. Correct If verification fails: backtrack to last known-good step, fix, recompute, re-verify No "doubling down" on wrong answers 7. Deliver Exact answer (not approximate), domain con

Chenrui Hu 2026-05-31 14:25 7 原文
AI 资讯 Dev.to

How Zone01 Kisumu "Build from Scratch" Approach Transformed Me from a Framework User to a Problem Solver

The Moment I Realized I Didn't Really Know JavaScript I was 2 months into learning JavaScript. I could use .map(), .filter(), .reduce() like any bootcamp grad. I felt confident. Then my instructor asked me one question: "How does .reverse() actually work?" I froze. I had used it hundreds of times. But I had no idea what was happening inside. I was a user, not a builder. That was the day everything changed. The 01EDU Difference: Build Tools, Not Just Use Them Most coding courses teach you to use built-in methods. 01EDU does something different. They disable the built-in methods. Then they say: "Now build it yourself." No .split(). No .join(). No .indexOf(). No .slice(). Just you, a text editor, and your brain. What I Built in 2 Weeks (Without Using Built-ins) Here are the JavaScript methods I re-created from scratch: Method What I Learned abs() Math is logic, not magic multiply(), divide(), modulo() Arithmetic is repeated addition/subtraction indexOf(), lastIndexOf(), includes() Searching is just looping and comparing slice() Negative indexes count from the end reverse() Arrays and strings are both indexed collections join() Building strings step by step split() Parsing is character-by-character inspection round(), floor(), ceil(), trunc() Decimals are just numbers between whole numbers Each function took hours of thinking, failing, debugging, and finally — understanding. The Most Painful Lesson: Loops The first time I tried to build repeat() without using .repeat(), I wrote an infinite loop. My computer froze. I had to force restart. That failure taught me more than any working code ever could. I learned to trace each iteration mentally. I learned to check my exit conditions. I learned to respect the loop. You don't truly understand loops until you've crashed your computer with one. What 01EDU Taught Me That No Bootcamp Could I learned how computers think, not just how to write code When you build .split() from scratch, you understand string parsing at a deep level.

Kevin Nambubbi 2026-05-31 14:24 8 原文