AI 资讯
Paper Reading Notes: [JEPA]
[Paper Notes] JEPA: Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture 🔗 TL;DR: JEPA learns a a generalized semantic representation with less data pairs by predicting missing information in the embedding space , which helps it disregard unnecessary noisy from input(pixel)-level details and learns at a higher abstraction level with good semantic generalization. 1. Innovation & Significance The Bottleneck: Image-text data pair labels are hard to find Pixel level pre-training paired & data augmentation are strongly biased towards trained data distribution, hard to determine proper generalization and level of abstraction. JEA's (Joint Embedding Architecture) collapse probelm: encoder & decoder attempts to cheat by always landing on trivial constant when predicting itself (reconstruction) and gets away with an easy Error=0. The Solution: > Chain-of-thought ⭕ Mask pre-training to reduce data & generalize↓❌ Bad/lower semantic representation without semantic target, could be learning noisy local pixel correlation↓⭕ Learn at the embedding level to omit pixel input and generalize⭕ Adds context encoder & positional encoding to inject context and force model to pick up image inherent structure from reconstructing multiple masked patches with one target.↓❌ JEAs wants to cheat: if I always map all pixels to a constant for both the predictor and end target encoder then the reconstruction error is always collapsed to zero! Hehe~ ↓ ⭕ EMA (Exponential moving avg.): Update target encoder parameters from the EMA of context encoders. This 'delays' the target encoder to prevent collapsing (a trick from the BYOL paper[2020], proven essential to training JEAs with ViT). 2. Model & High-Level Intuitions 2.1 Model Architecture Input: randomly samples block masks from original image within certain aspect ratio changes, and apply mask for context image 2.1.2 Context Context Encoder: ViT encodes context image to embedding SxS_x S x Mask Token : an [1,D] random
AI 资讯
I built an AI conversation simulator because I kept chickening out of real talks
Last year I needed to ask for a raise. I knew my number, I'd read the guides, I had bullet points in my notes app. Then my manager said "let's chat about your goals for next quarter" and I said "sounds great, looking forward to it" and hung up. Never brought up money. Same thing kept happening elsewhere. Coworker taking credit for my work, I said nothing. Relationship that should've ended months earlier, I kept postponing. I always knew what to say. I just couldn't say it with someone actually looking at me. So I started building a thing to practice on. That thing became cosskill . What it actually is You pick a persona, tell it the situation in a sentence, and start talking. The persona doesn't help you. It holds position and pushes back. You practice not folding. Think of it as a flight simulator for hard conversations. You rehearse until your opener comes out steady, then go do the real thing. 20 personas across five categories: Operators (Musk, Jobs): first-principles thinking, harsh product feedback Strategists (Trump, Buffett): treat everything as a deal or a bet Relationship (Ex, Coworker): breakups, workplace friction, family money Philosophy (Socrates, Aurelius, Confucius, Sun Tzu, four more): each tradition frames problems differently Psychology (Rogers, Rosenberg, Ellis, Frankl, Kahneman, Jung): therapeutic frameworks on real situations These aren't celebrity impressions. The Buffett persona won't hype your startup idea. It'll ask "what's the downside?" and keep asking until you have something concrete. Tech stack Next.js 16 on Cloudflare Workers. DeepSeek for inference. Cloudflare D1 (SQLite at edge) for the bits that need to persist. No user accounts, chat history lives in localStorage. Monthly cost stays low enough that the free tier (10 messages/day) doesn't worry me. Why I made these choices DeepSeek instead of GPT-4/Claude. Each conversation is 10-30 messages. At GPT-4 pricing a free product bleeds money. DeepSeek gives maybe 90% of the quality for
AI 资讯
From Axios to alova: how we cut 80 lines to 5
Frontend request code often involves repetitive state management. This article compares Axios and alova through a paginated list example, analyzing how request strategization reduces boilerplate and when it's a good fit. The Pattern: Paginated List in Two Ways A common requirement: fetch a user list with pagination. Approach 1: Axios const [ data , setData ] = useState ([]); const [ page , setPage ] = useState ( 1 ); const [ total , setTotal ] = useState ( 0 ); const [ loading , setLoading ] = useState ( false ); const [ error , setError ] = useState ( null ); const fetchUsers = async ( currentPage ) => { setLoading ( true ); setError ( null ); try { const res = await axios . get ( ' /api/users ' , { params : { page : currentPage , pageSize : 10 }, }); setData ( res . data . list ); setTotal ( res . data . total ); } catch ( e ) { setError ( e . message ); } finally { setLoading ( false ); } }; useEffect (() => { fetchUsers ( page ); }, [ page ]); This pattern appears in nearly every data-fetching component. The actual business logic — GET /api/users — occupies a single line. The rest is infrastructure: state declarations, loading toggles, error handling, and effect management. Approach 2: alova with usePagination const { data , total , loading , error , page , pageSize , nextPage , prevPage , } = usePagination ( ( page , pageSize ) => alovaInstance . Get ( ' /api/users ' , { params : { page , pageSize }, }), { page : 1 , pageSize : 10 } ); Both implementations are functionally identical. The key difference is where the state management logic lives: in the component (Axios) vs. inside the hook (alova). What Changed Component of Axios version Handled by alova loading state + toggling Managed internally by usePagination error state + try/catch Managed internally by usePagination data state + assignment Returned as reactive value page state + change handler Built-in nextPage / prevPage total state extraction Extracted from response automatically useEffect dependency tr
AI 资讯
Agentic Web3: Automating Blockchain Workflows with Hermes
This is a submission for the Hermes Agent Challenge Agentic Web3: Automating Blockchain Workflows with Hermes Tags: #hermesagentchallenge , #web3 , #agents , #solana The blockchain industry has spent the last decade building decentralized, permissionless infrastructure. However, the user experience layer interacting with this infrastructure remains overwhelmingly manual. Decentralized applications (dApps) require users to constantly monitor markets, parse complex data, and manually sign every transaction. The next evolution of Web3 isn't just about faster blockchains; it is about autonomous execution. By integrating large language models and agentic frameworks with smart contracts, we can transition from a paradigm of manual execution to intent-based autonomy . In this article, we will explore how to bridge the gap between AI and decentralized networks by automating blockchain workflows using the Hermes Agent framework. We will look at the architecture of an on-chain agent, how it reads and writes to a network, and how high-performance environments like Solana are making these agentic experiences viable. The Paradigm Shift: From Passive Wallets to Active Agents Currently, most AI in Web3 is limited to read-only analytical tools—chatbots that can summarize a smart contract or pull token prices from an API. While useful, these are fundamentally passive systems. An active agent is different. Powered by a framework like Hermes Agent, an active agent can: Observe: Continuously monitor on-chain events via RPC nodes or webhooks. Reason: Use its LLM core to interpret those events against a set of user-defined goals or risk parameters. Act: Formulate a transaction, sign it via a secure wallet environment, and broadcast it to the network. This opens up massive possibilities. Imagine an agent that automatically manages your decentralized finance (DeFi) positions, rebalancing a portfolio based on yield changes across different protocols. Or consider fully on-chain gaming, where
开源项目
How do you batch optimize images for your projects?
I’m asking since I’m working on an all-in-one app for bulk image optimizing and would like to know more about my potential users. What is your workflow is like today? What tools and processes do you have? I appreciate any and all info you are open to share. Thanks! submitted by /u/TrapShot7 [link] [留言]
AI 资讯
How I Rebuilt My Entire User Feedback Workflow with FeedLog (And Why I Ditched Canny)
Six months into running my SaaS, my "feedback system" was three browser tabs, a starred Gmail folder, and a sticky note on my monitor that said "check Discord." That was the whole system. It held together until the day I found a three-paragraph email from a paying user — a genuinely detailed feature request with a real use case — sitting unread for 24 days. His last line was: "Happy to pay more if you can support this." I replied the same afternoon I found it. His reply: "Switched last week, thanks anyway." That was the moment I stopped treating feedback management as a nice-to-have. Why the usual fixes didn't fix anything I tried the obvious things first. I want to document them because I see a lot of people cycling through the same failed solutions. Notion database 🪦 Built a beautiful one. Color-coded tags, priority columns, status tracking. It lasted 11 days before nobody — including me — was maintaining it. The friction of "open Notion, find the right database, fill in six fields" is invisible when you're designing the system and fatal when you're in the middle of a support conversation. Airtable form 🪦 Better entry point, still disconnected from where users actually were when they had feedback. Nobody bookmarks your Airtable form. They DM you on Discord and you think "I'll add that later" and you don't. Canny — this one actually worked, for a while I genuinely liked Canny. Clean interface, users could upvote requests, I could see what was popular. It felt like a real system. Then our user count grew and the pricing tier jumped. I was looking at $99/month for a feedback board for a product still finding its footing. That's not a moral judgment on Canny — it's a fair product — but for a bootstrapped indie dev, it started feeling like a tax on momentum. The deeper problem with all three solutions was the same: they were inboxes, not loops. User submits → enters the void → user never knows if anyone saw it → user assumes nobody did → trust erodes → churn. I had bui
AI 资讯
I'm looking for someone to start a project with me.
Hi, I'm Brazilian and I'm a web developer. I'm eager to start a project (without involving money at the moment). I'd like to start developing something together. I've been working in this field for a while now, I have several projects underway, and I'd like to start another one with someone.If anyone is interested, I can give more details about myself and some ideas I've had in private 🙂 I think two people working together learn more and wouldn't spend anything; it would only involve money for selling the website submitted by /u/BitesBitesou [link] [留言]
开源项目
From Abandoned Side Project to Full Wellness Hub — My GitHub Copilot Glow Up ✨
``This is a submission for the GitHub Finish-Up-A-Thon Challenge What I...
AI 资讯
Production-Ready Logging: An Agnostic ELK Stack Setup for Node.js (with a 512MB RAM Local Constraint)
The Logging Nightmare Deploying microservices across Multi-Cloud environments using tools like Terraform is an exhilarating milestone. But the moment something breaks, that excitement quickly turns into a nightmare. The SSH Grind : If you find yourself SSH-ing into disparate instances just to run tail -f and grep through scattered log files, you're doing it wrong. The Agnostic Approach : The industry standard demands Centralized Logging, but chaining your application to vendor-specific solutions like AWS CloudWatch or GCP Cloud Logging limits your architectural freedom. Implementing a true "Cloud-Agnostic" ELK stack gives you back control over your observability data. Clean Architecture & The Non-Blocking Logger Factory Building this robust observability pipeline requires adhering to Clean Architecture principles, specifically through a Non-Blocking Logger Factory. Standardized Interface : By wrapping modern logging libraries like Winston or Pino , we standardize our application's logging interface. The Secret Sauce : The winston-elasticsearch transport module buffers your logs and pushes them directly to your Elasticsearch cluster in the background. Non-Blocking : This architectural choice is crucial: it ensures that high-volume log streaming happens without blocking the Node.js event loop . Here is how the data flows through the system: Resilience Fallback (The Failsafe) A centralized system introduces a dangerous dependency. Your logging infrastructure must never be the reason your application crashes. The Threat : If the remote Elasticsearch cluster is unreachable due to network partitions or rate limits, a poorly configured logger will throw uncaught exceptions, bringing down the app. The Solution : We implement a strict Resilience Fallback (Failsafe) mechanism. The transport module safely catches the connection errors and seamlessly falls back to standard output (console), guaranteeing continuous operation. The 512MB Local-Test Challenge While this setup is a
AI 资讯
I Rebuilt My Karaoke App So Everyone's Phone Could Be a Remote
This is a submission for the GitHub Finish-Up-A-Thon Challenge What I Built VKara is a browser-based karaoke room app for singing at home with friends or family. It is not trying to replace YouTube. YouTube is already great at playing videos. It already has almost every karaoke song we need. But YouTube is not really designed to manage a karaoke night where many people want to choose songs together. That is the gap VKara tries to fill. You open VKara on a TV or laptop as the main playback screen. Everyone else joins the same room from their phone using a 4-digit room code or QR code. Then anyone can search for songs, add them to the queue, pause, resume, or skip. The TV only needs to play the video. Everyone's phone becomes their own remote. That is the whole idea. Simple enough to explain in one sentence. Not simple enough to build in one weekend. I learned that part the hard way. Demo Links: Live demo: https://vkara.vercel.app/en GitHub repo: https://github.com/lehuygiang28/vkara Before branch: https://github.com/lehuygiang28/vkara/tree/before Old backend repo: https://github.com/lehuygiang28/vkara-api Small warning: the demo is running on limited resources, so if it is slow, please give it a moment. My wallet is still a student wallet. lol. The flow is: Open VKara on a TV or laptop. Join the room from a phone by code or QR. Search for a karaoke video. Add it to the shared queue. Control playback together. Before: the idea worked, but the product still felt like a video app squeezed into a karaoke use case. After: the mobile flow is now focused on joining, searching, choosing an action, and controlling playback. The Comeback Story I started VKara around early 2025. At that time, my goal was very personal. I wanted a better way to sing karaoke at home with friends. The normal setup was: open YouTube on a TV, search for karaoke videos, and pass control around. It worked, but it was awkward. One person was searching. Another person accidentally played a video immedia
AI 资讯
public-apis: what 438k stars actually buy you, and what they don't
Repository: public-apis/public-apis What public-apis actually is public-apis is a community-curated directory of free and public APIs, maintained by contributors together with staff at APILayer. It is not a library, SDK, or gateway: there is no package to import and nothing to run in production. The repository is essentially one very large, structured README that catalogs APIs across roughly fifty categories, from Animals and Anime to Finance, Machine Learning, Security, and Weather. Each entry is a row in a table with five columns: the API, a short description, the authentication model ( apiKey , OAuth , or none), whether it serves over HTTPS, and whether it sets permissive CORS headers. That last detail is the part most engineers undervalue. Why engineers keep coming back to it The star count, now past 438,000, is less interesting than the metadata discipline. When you are prototyping and need a currency-exchange or geocoding endpoint, the Auth/HTTPS/CORS columns let you filter candidates before you ever open a browser tab. "No auth, HTTPS yes, CORS yes" tells you that you can call the endpoint directly from a front-end spike without standing up a proxy or registering for a key. For throwaway demos, hackathons, internal tools, and teaching material, that triage saves real time. The category index doubles as a map of what kinds of public data are actually available, which helps when you are scoping whether an idea is even feasible. How it is maintained Curation is manual and community-driven: changes arrive as pull requests against the README, governed by a contributing guide, with issues and PRs as the moderation surface. The project's primary language is Python, reflecting validation tooling that checks entries rather than any runtime you would consume. There is also a separate companion project that exposes the list itself as an API. The model is simple and has clearly scaled, but "manually curated" is both the strength and the weakness. Limitations worth statin
开发者
Best way to build android app without learning java?
So i've been putting off this project for months but i really need to bu͏ild android a͏pp for this side business idea i have. Problem is i don't know java and honestly don't have time to learn it properly right now. I've heard there are ways to build apps without actually coding everything from scratch but not sure what's legit vs what's just marketing bs. I can handle basic web stuff but mobile development seems like a whole different beast. Anyone here successfully built an android app without going the traditional java route? What tools or platforms actually work and don't make your app look like garbage? submitted by /u/Significant_Law5994 [link] [留言]
AI 资讯
unpopular opinion: most ecommerce companies don't need headless or composable
If you're a fashion or DTC brand doing under €50M GMV, you'll get more out of a monolithic platform than a composable build, because the complexity tax for that flexibility runs higher than a year of performance marketing for most brands at that tier. Composable was sold as freedom, but for most brands at small-to-mid GMV it lands as engineering-org overhead you fund forever just to keep the lights on, while monolithic platforms ship 80% of what most brands need with one team and one contract. The exception is brands running multi-brand groups across countries or genuine enterprise peaks, where composable does pay off, but you want the platforms built by retailers rather than the ones built for analyst slides. Centra, NewStore, and SCAYLE land closer to that brief than commercetools or open-source MACH stacks. If a vendor opens their pitch with microservices and best-of-breed architecture, they're selling you an engineering project, and most brands just need a platform that ships things. submitted by /u/Majestic_Shoulder188 [link] [留言]
科技前沿
Everyone Has Their Targets Set on the MacBook Neo
Dell, Microsoft, and others are unveiling new laptops to compete directly with the Neo, but not all are learning the right lessons from Apple.
开发者
When Duplicate Code Is the Better Design
You've seen the developer. Maybe you are the developer. They discover DRY — ✨ Don't Repeat Yourself...
开发者
🌐OS May Recap: Learning to Navigate the Open-Source Galactica
In May, I continued my "One Commit a Day" Challenge and spent more time contributing across different open-source projects. Compared to April, I was able to contribute a bit more and explore a wider variety of repositories. Repositories That Stood Out Some of the projects that left the biggest impression on me were: python-odpt Huggin Face Context Course Human Signal ML ScribeSVG A Stable Checkpoint One milestone I was happy about this month was reaching a stable checkpoint for my Tokyo MCP Server project. It is still a work in progress, but getting to a point where the project feels stable enough to build upon was a satisfying moment. Documentation Matters Another contribution that stood out was helping improve a python-odpt README documentation . It wasn't a large technical contribution, but it reminded me that making a project easier for others to understand can be just as valuable as writing code. Good documentation lowers the barrier for future contributors. Sometimes, a clearer README can help more people than a small code change. Learning Beyond Python One practical lesson I learned this month was that being a Python-focused contributor doesn't mean I can ignore the JavaScript ecosystem . While working with different repositories, I finally installed Node.js and started using npm . Many modern open-source projects rely on TypeScript-based tooling, build systems, or development workflows, and understanding those tools makes contributing much easier. The Biggest Challenge: Finding Information And Communication Matters The biggest challenge I faced wasn't coding. It was documentation. Every repository has its own way of organizing information. There are definitely common patterns, but every project also develops its own style over time. Sometimes the information I need is in the README. Sometimes it's in a wiki. Sometimes it's buried in a docs folder several levels deep. And sometimes it's spread across all three. Open Source Is Also About Navigation As a contri
AI 资讯
I built PhysioFlow — clinic software for Indian physiotherapists, solo in a week
A physiotherapist asked me a simple question a couple of months ago: "Can you build something to run my whole clinic?" So I did — solo, in about a week. Here's the full 2.5-minute walkthrough 👇 What PhysioFlow does PhysioFlow runs an entire physiotherapy clinic from one screen — built for India (₹, GST, WhatsApp, +91): Dashboard — attendance, collections & pending bookings at a glance Patient files — recharge session packs, track usage, auto-generate GST-ready bills Attendance in seconds with a QR scan Online bookings that convert straight into a patient file Reports — daily ledger, revenue, CSV/PDF export Patient portal — patients see their own sessions & prescriptions The stack Next.js + Supabase + TypeScript — multi-tenant, with row-level security so no clinic's data ever leaks to another. Try it Live now with a 14-day free trial, no card needed → https://physioflow.devfrend.com I'm Amar, a full-stack & AI engineer. I design and build products like this end-to-end. I'm open to work — SaaS builds, MCP servers, LLM apps & automation. Reach me on LinkedIn .
AI 资讯
Before I Would Trust an Agent's Memory, I Would Audit Its Authority
This is a submission for the Hermes Agent Challenge , under the Write About Hermes Agent prompt. I've spent the last week testing AI memory failure modes in a public evaluation harness. That work changed how I read agent memory systems. This is a writing submission, not a build submission. I did not build a Hermes Agent project for this challenge. I am writing from the perspective of someone testing how memory failures show up once agents can act. So when I look at Hermes Agent, the question I care about is not only: Can the agent remember useful things? The harder question is: When memory conflicts, which memory is allowed to govern the agent's action? That distinction matters. Hermes Agent is interesting because it is not just a chat interface. Its documentation describes an open-source agentic system with tool use, project context, persistent memory, skills, browser automation, checkpoints, delegation, scheduled tasks, and multiple memory providers. That is exactly the kind of system where memory stops being a convenience feature and starts becoming part of the agent's operating boundary. If an agent can run tools, edit files, browse, delegate work, schedule tasks, and remember across sessions, then memory is no longer just "context." Memory becomes governance. The Memory Problem I Would Watch For In a simple chatbot, bad memory is annoying. In an agent, bad memory can become operational. The failure mode is not only that the agent forgets something. Sometimes the more dangerous failure is that it remembers the wrong thing too confidently. A memory can be: relevant but stale, relevant but low-authority, relevant but superseded, relevant but only context, relevant but not allowed to determine the action. That is the distinction my own tests kept running into. Retrieval systems are usually good at answering: What memory is closest to the user's request? But safety often depends on a different question: What memory is allowed to decide what the agent should do? Thos
AI 资讯
Building Hermes Financial Agent: An Explainable AI Copilot for EGX Investors
Overview I built Hermes Financial Agent — an AI-powered financial assistant for investors in the Egyptian Exchange (EGX). The goal: not just calculate portfolio value, but explain risks in a transparent, auditable way. Features Real EGX market data via Yahoo Finance Portfolio tracking and valuation Daily financial reports Telegram-user portfolio isolation Cached quote fallback for resilience Explainable risk insights Explainable Risk Intelligence Unlike traditional trackers, Hermes surfaces: Portfolio concentration risk Stale market data exposure Quote coverage percentage Valuation gaps from unavailable data Users understand confidence level behind their valuation — not just the number. Technical Stack Hermes Agent framework Python GitHub Actions (automated smoke testing) Offline-safe quote fallback architecture Repository 🔗 GitHub Repository Future Roadmap News sentiment analysis Investment thesis tracking Market briefing generation Advanced financial reasoning agents hermeschallenge
AI 资讯
Handling large images & files in a real time chat application
Hey guys, would love to get some feedback on whether my approach here makes sense. I’m building a real-time chat application where users can upload and receive images/files. Files can be fairly large (up to ~100MB). Current stack is: fastapi, websockets, tanstack, postgres and redis. I use GCP as my cloud provider. My current flow is: Backend generates a signed URL Frontend uploads directly to a GCS bucket A Cloud Function handles post-upload processing For downloads, the frontend fetches directly from GCS using redirects/signed URLs so the backend doesn’t become a bottleneck This architecture works great for smaller files (<30MB), but once I started testing larger uploads (100MB images/videos), I noticed very high memory consumption during processing (btw pagination & virtualization is used throughout the project). I’m trying to figure out what’s considered best practice here for large media uploads in chat systems: Should compression/downscaling happen client-side or server-side (currently there is not compression at all)? Also, Is it common to generate thumbnails/previews (for images) separately while keeping the original untouched? Should I stream uploads instead of buffering them? Are Cloud Functions even the right choice for heavy file processing? For images specifically, I’m considering: client-side compression before upload, automatic thumbnail generation, storing multiple resolutions, converting to formats like WebP/AVIF. Would love to hear how you guys handle this in production systems. Thanks! submitted by /u/omry8880 [link] [留言]