今日已更新 287 条资讯 | 累计 37192 条内容
关于我们

今日精选

HOT

最新资讯

共 37192 篇
第 1739/1860 页
AI 资讯 Dev.to

How RAGScope Knows Which Chunks Your LLM Actually Used

How RAGScope Knows Which Chunks Your LLM Actually Used Your retriever fetched 10 chunks. Your LLM only used 3. RAGScope shows a precision score of 30 out of 100. The question every new user asks: how does it know? There is no OpenTelemetry attribute that says "this chunk was in the context window." RAGScope infers it — and the way it does this is the most consequential piece of engineering in the whole tool. There Is No "In Context" Attribute in OTel The OpenTelemetry semantic conventions for generative AI ( gen_ai.* ) define attributes for model, input/output tokens, and retrieved documents. They do not define anything like gen_ai.chunk.reached_llm or gen_ai.retrieval.used_document_ids . When your RETRIEVER span fires, you get a list of documents. When your LLM span fires, you get a prompt and a completion. The two spans are connected by a parent-child trace relationship — but there is no attribute that maps which retrieved documents appear in which prompt. This gap matters. A reranker might drop 7 of your 10 chunks. Your application code might apply a token budget and truncate 4 more. From the trace alone, you cannot tell. RAGScope needs this information to compute the precision sub-score — the highest-weighted metric at 40% of the overall score. Getting it wrong would make precision meaningless. The Substring Match — How assembleContext Works RAGScope's answer is in src/enrichment/pipeline.ts , in a function called assembleContext : function assembleContext ( chunks : RagChunk [], llmSpans : ParsedSpan []): RagChunk [] { const llmPrompts = llmSpans . map (( s ) => s . prompt ). filter (( p ): p is string => !! p ); if ( llmPrompts . length === 0 ) return chunks ; let position = 0 ; return chunks . map (( chunk ) => { if ( ! chunk . content ) return chunk ; const inContext = llmPrompts . some (( p ) => p . includes ( chunk . content ! )); if ( inContext ) { return { ... chunk , inContext : true , contextPosition : position ++ }; } return { ... chunk , inContext :

Siddharth Pandey 2026-05-31 17:46 10 原文
AI 资讯 Dev.to

Arazzo Visualizer: Run API Workflows in VS Code

Most apps don't just call one API endpoint. They call a whole chain of them. For example, you might log in, get a token, and then pass that token to another service. Tracking these multi-step chains can get messy quickly. To help fix this, the OpenAPI Initiative created the Arazzo Specification . It gives us a standard way to link different endpoints into clear workflows. But writing these workflow files by hand in a regular text editor is tough. It is very easy to lose track of how data moves from one step to the next. That is why I built Arazzo Visualizer for VS Code. It is a free, open-source extension that makes the Arazzo spec visual and easy to use. Live Interactive Graphs The extension reads your workflow files and turns them into interactive maps on the fly. See Data Flow: Look at exactly how data moves between steps. Catch Errors Early: Spot broken paths before you even run your code. Clean Layouts: Navigate large workflows without getting lost in thousands of lines of text. Built-In Workflow Runner Seeing the map is great, but testing it is even better. The tool has a step-by-step runner built right into your editor. Run a single step or execute the whole chain. See real-time data payloads and HTTP headers. Watch requests happen live to pinpoint bugs fast. Give it a Try The project is fully open source, and you can grab it or check out the code using the links below: Download: Install it directly from the VS Code Marketplace . Source Code: Check out the repository, report bugs, or contribute on GitHub . Deep Dive: Read my full technical breakdown and design on Medium . If you are working with API chains, I would love for you to try it out. Drop your feedback in the comments below! Note: Arazzo v1.1.0 is out with official AsyncAPI support. I am currently updating the VS Code extension to support these new features. Stay tuned for future updates!

Himeth Walgampaya 2026-05-31 17:43 15 原文
AI 资讯 Dev.to

Developer will need to understand lambda by 2026

I used to deploy Node.js apps on EC2 and manage servers like it was my second job. Port configs. PM2 restarts. Nginx rewrites. SSL renewals. Then I ran my first AWS Lambda function. 80% of that work is gone. Here's what Lambda actually does that nobody explains clearly: → You write a function → AWS runs it ONLY when triggered → You pay for milliseconds of execution → It scales from 1 to 1,000,000 requests without you touching anything As a full-stack developer in Bahrain, preparing for my AWS Developer Associate exam, this is the shift that changes how you think about backend architecture. Not "how do I manage a server" but "what should happen when this event fires." That mental model switch took me a week to fully get. I'm documenting everything as I study. Drop a 🔥 if you want me to share my Lambda notes weekly.

Hamza Rafique 2026-05-31 17:38 13 原文
AI 资讯 Dev.to

App Size: A Battle for Every Kilobyte, or Prioritizing Functionality?

Throughout my software development career, especially over the last 20 years, I've constantly sought a balance between system and software aspects. One of these balancing points is the issue of application size. On one side, I have an engineering instinct that says, "Every kilobyte matters, every excess is a cost," while on the other, a pragmatic perspective that argues, "Delivering value-adding functionality to the user as quickly as possible is essential." Navigating between these two extremes has been a significant part of my career. So, should we really fight for every kilobyte? Or should we prioritize functionality? This is a gray area that varies depending on the project's context, target audience, and even deployment model. In my experience, there's no clear-cut answer to this question, but the best results can be achieved by asking the right questions and making conscious trade-offs. In this post, I'll share my thoughts and observations from different fields on this topic. The Harsh Realities of Size in the Mobile World The importance of size in mobile applications is undeniable. Especially in markets like Turkey, where mobile internet is widespread but not always fast or unlimited, app size directly affects download rates, user experience, and even how long an app stays on a device. Years ago, while developing my own Android spam application, I experienced this reality repeatedly. The smaller I made the app, the higher the download and installation rates became. For example, at one point, my app's main package was around 12MB. After several rounds of optimization (cleaning up unnecessary resources, code compression with ProGuard, optimizing native libraries), I managed to reduce this size to 7MB. This 40% reduction directly reflected in my Play Store download statistics. This difference was critical, especially for users with low-bandwidth connections or those using capped mobile data. Some users even provided feedback stating they refrained from downloadin

Mustafa ERBAY 2026-05-31 17:36 10 原文
AI 资讯 Dev.to

How AI reads your website, and what that means for the people who build it

By Takeshi Yokoyama — Onecarat Labs Hi. I'm Yokoyama, and I build a local-first AI text editor as a side project, along with a few other experimental tools. Working on them, I keep running into the same question about where the web is going. This post is one observation, plus a small experiment I built to test it — including a Chrome extension you can actually try. The short version: I think websites will increasingly be read through AI agents, reshaped per reader, on the fly. And once that happens, there's a clear gap between sites that are easy for an AI to read and sites that aren't. What's starting to happen Until now, people read websites as websites. You open the top page, follow the menu, read the body, click a button — tracing the path the maker designed. As local AI and AI agents become normal, that breaks. People stop opening the page directly. They tell an AI what they want — "Can I try this quickly?" , "I just want to check it's safe" , "Just the gist" — and the AI reads the web and reshapes it into the form that reader wants. What the reader receives is no longer the layout the maker built. This isn't speculation. The idea that AI generates the interface for the reader already has a name — Generative UI — and it's one of the hottest areas in frontend right now, with Google, Vercel and others building toward it. But notice who's holding the pen in almost every version of that story: the site , or an AI embedded in an app — something under the maker's control. What I'm looking at is one step past that: a local AI, in the reader's own hands, reshaping any site into that person's preferred form — with no involvement from the maker at all. The initiative moves from the maker to the reader. The part that nags at me as a builder I build software too. So this shift nags at me. A site carries its maker's intent and rights. The order things appear in, what gets emphasized, the tone. Design, copy, flow — all of it is deliberate. Having an AI quietly reorder, rewri

Onecarat Labs/Takeshi Yokoyama 2026-05-31 17:33 14 原文
AI 资讯 Dev.to

How I Built a Live Football Platform That Doesn't Fall Apart Under Load

A walkthrough of the architecture decisions behind Flacron Gamezone a production full-stack app built with Next.js, Express, PostgreSQL, and Redis. When a client approached me to build a live football match discovery platform, the requirements sounded straightforward on the surface: show live scores, let users subscribe, handle authentication. But the moment you start thinking about how those pieces connect in production, straightforward gets complicated fast. This is the story of how I designed the backend for Flacron Gamezone — what decisions I made, why I made them, and what broke along the way. Table of Contents The Problem With "Just Building It" The Architecture: Four Distinct Layers Why This Matters to a Client The Bug That Taught Me Something Real The Full Stack at a Glance What I'd Do Differently The Problem With "Just Building It" The easiest version of this app is a single Express file: one route handler that queries the database, formats the data, and sends a response. I've seen this pattern in tutorials everywhere. It works for demos. It falls apart in production. The problems are predictable: you can't test business logic without hitting the database, a change in one feature quietly breaks another, and the moment a second developer joins the codebase, nobody knows where anything lives. I wanted to build something I could actually be proud to show an employer or a client. That meant committing to a proper layered architecture from day one, even on a project this size. The Architecture: Four Distinct Layers The entire Express backend is organized into four layers. Each layer has one job and talks only to the layer directly below it. Route → Controller → Service → Repository Here's what each one actually does. Routes are just maps. They declare that POST /api/v1/subscriptions exists, attach the auth middleware, and hand off to the controller. No logic lives here. Controllers handle the HTTP boundary. They extract data from req.body or req.params , call th

Ahmed Ali 2026-05-31 17:33 8 原文
AI 资讯 Reddit r/MachineLearning

I built mlx-Chronos — a community benchmark leaderboard for local LLM engines on Apple Silicon (oMLX, Rapid-MLX, mlx-lm, Ollama) [P]

Hey! I'm a CS student and I got tired of not being able to compare MLX inference engines properly — every benchmark out there is either made by the engine's own developers, runs on an M3 Ultra nobody has, or just shows tok/s with zero context. So I built mlx-Chronos — a small open source CLI tool that runs a standardized benchmark protocol on your Mac and lets you submit your results to a shared community leaderboard. What it measures: Cold and cached TTFT (Time to First Token), with a proper methodology — unique prompts per trial, cache priming, no interleaved phases Throughput (tok/s), with mean/stddev/min/max across repeated trials Engine process RSS and system RAM peak, sampled continuously during inference Thermal state and hardware info Supported engines: oMLX, Rapid-MLX, mlx-lm, Ollama (MLX backend) The leaderboard is basically empty right now since I only have an M2 8GB. Would love results from M3 Max, M4, M4 Ultra, or anything with more RAM — that's where things get actually interesting. → Leaderboard: https://igurss.github.io/mlx-chronos → GitHub: https://github.com/igurss/mlx-chronos → Install: pip install mlx-chronos It's early, the methodology is documented (there's a methodology.md if you want to pick it apart), and I'm 100% open to feedback, contributions, and getting told what I'm doing wrong. The goal is just to have one place where you can compare engines on your specific hardware instead of trusting someone else's numbers. submitted by /u/igor__004 [link] [留言]

/u/igor__004 2026-05-31 16:26 6 原文
AI 资讯 Reddit r/webdev

Have DIY website builders had an effect on the web-dev industry?

I've recently been getting reacquainted with frontend dev after having spent the last few years focusing mostly on design. I plan to do a combination of both again from now on. The last full-time job I was at we used drag-and-drop website builders that seem to be all the rage these days. I gotta say it was a truly painful experience. The websites we would handover would be a complete mess, clients would have to edit columns, widgets, layouts themselves which just felt like a recipe for disaster. We would often have to duplicate various widgets such as image sliders, one of mobile and one for desktop. I'd then have to explain to clients that they needed to update them twice, with the same content. Also, a lot of times I'd have to set up HTML widgets because of how limited the software was, so I'd also have to explain to clients that they'd sometimes need to edit actual code when simply updating content on their site. But the concerning thing is, I've had a number of other job interviews at marketing agencies lately and it seems like EVERYONE is doing the same thing. I also sometimes hear on podcasts people describing default WordPress as being 'old fashioned', as though the future of web dev is just doing everything with these hideous website builders WTF 🤮 Are these people mainly just taking the basic / shitty web dev projects while leaving the proper ones to those who know how to do it properly? or have a lot of developers lost out on work because of these? submitted by /u/Weekly_Frosting_5868 [link] [留言]

/u/Weekly_Frosting_5868 2026-05-31 16:07 5 原文
AI 资讯 Smashing Magazine

June Is For Exploring (2026 Wallpapers Edition)

Let’s kick off June — and the beginning of summer — with some fresh inspiration! Artists and designers from across the globe once again tickled their creativity to welcome the new month with a new collection of desktop wallpapers. Enjoy!

hello@smashingmagazine.com (Cosima Mielke) 2026-05-31 16:00 4 原文