今日已更新 133 条资讯 | 累计 37382 条内容
关于我们

标签:#huggingface

找到 7 篇相关文章

AI 资讯

GGUF vs GPTQ vs AWQ: Which Quantization Format Should You Actually Use?

Running open-source Large Language Models (LLMs) used to be a luxury reserved for developers with enterprise-grade server rooms. If you didn't have dual A100 GPUs sitting under your desk, running a modern 8B or 14B parameter model was a one-way ticket to Out-Of-Memory (OOM) crashes and frozen systems. Then came quantization. By compressing 16-bit floating-point weights (FP16) down to 4-bit or 8-bit integers, quantization slashes the VRAM footprint of LLMs by 70% or more, often with barely noticeable drops in accuracy. But as you browse Hugging Face for a model, you are immediately hit with a wall of acronyms: GGUF, GPTQ, and AWQ. Which format actually fits your hardware? Which one delivers the fastest tokens-per-second? And how do you generate these files without melting your local machine? Let's break down the definitive differences so you can choose the exact format your pipeline needs. 1. GGUF: The King of Local Hardware and CPU Offloading Developed by the team behind llama.cpp, GGUF (GPT-Generated Unified Format) completely revolutionized local LLM execution. How it works: Traditional formats require a powerful GPU to load a model. GGUF changes the rules by allowing CPU offloading. If a model requires 12 GB of VRAM but your graphics card only has 8 GB, GGUF splits the layers: it loads 8 GB into your GPU and shunts the remaining 4 GB to your system RAM and CPU. The trade-off: While running models on system RAM is significantly slower than running them purely on a graphics card, GGUF ensures the model actually runs. It turns a guaranteed system crash into a functional, runnable local AI. If you have a powerful GPU, GGUF can also run 100% on the graphics card for blistering speeds. Hardware: Apple Silicon MacBooks (M1/M2/M3), laptops with consumer Nvidia cards (e.g., RTX 3060/4060), or setups without a dedicated GPU. Use Case: Local application development, hobbyist exploration, and offline edge computing. 2. GPTQ: Enterprise-Grade Speed for Pure GPU Pipelines GPTQ

2026-08-09 原文 →
AI 资讯

Top AI Papers on Hugging Face - 2026-08-08

10 paper AI nổi bật nhất hôm nay trên Hugging Face: Agentic RL, computer-use, 3D world generation và hơn thế nữa Hôm nay, bảng xếp hạng paper trên Hugging Face cho thấy một xu hướng rất rõ: AI đang dịch chuyển từ mô hình “trả lời câu hỏi” sang hệ thống “thực hiện nhiệm vụ dài hơi” . Nhiều paper nổi bật tập trung vào agent, long-horizon planning, reward modeling, temporal reasoning, và khả năng hiểu không gian–thời gian trong môi trường phức tạp. Dưới đây là phần tổng hợp 10 paper được upvote cao nhất, với 4 góc nhìn cho mỗi bài: Bài toán Ý tưởng Điểm mới Ứng dụng thực tế 1) Recursive Synthesis for Long-Horizon Terminal Tasks Paper: 2608.05466 Project: Link Bài toán Nhiều tác vụ agent ngoài đời thực chỉ cho phản hồi ở cuối hành trình : làm xong một quy trình nhiều bước mới biết thành công hay thất bại. Đây là bài toán rất khó cho học tăng cường hoặc lập kế hoạch, vì tín hiệu thưởng quá thưa và không chỉ rõ lỗi nằm ở bước nào. Ý tưởng Paper này đề xuất hướng recursive synthesis : thay vì cố giải toàn bộ nhiệm vụ dài trong một lần, hệ thống chia bài toán thành các mục tiêu con, tổng hợp nghiệm từng phần, rồi xác minh và ghép lại theo cách đệ quy. Nói đơn giản, agent không “nhảy” từ đầu đến đích, mà xây một cây giải pháp: chia tác vụ lớn thành các tác vụ con, giải từng tác vụ con, kiểm chứng tính đúng đắn, hợp nhất thành nghiệm cuối. Điểm mới Điểm đáng chú ý là kết hợp giữa synthesis và verification cho các nhiệm vụ dài hơi. Khác với nhiều cách học agent chỉ dựa vào rollout và reward, hướng này nhấn mạnh tính đúng đắn có thể kiểm tra được , rất quan trọng khi xử lý terminal tasks. Ứng dụng thực tế Tự động hóa tác vụ doanh nghiệp nhiều bước AI thao tác phần mềm với quy trình dài Lập kế hoạch robot cần hoàn thành trọn vẹn nhiệm vụ Agent coding/workflow nơi chỉ bài test cuối cùng quyết định thành bại 2) AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Paper: 2608.05987 GitHub: Link Bài toán Agentic RL thường gặp hai vấn đề: chi phí khám phá cao và

2026-08-08 原文 →
AI 资讯

Top AI Papers on Hugging Face - 2026-07-30

10 paper AI nổi bật nhất trên Hugging Face hôm nay: robot thời gian thực, agentic search, coding agents và học tăng cường thế hệ mới Hôm nay, top paper được upvote cao trên Hugging Face cho thấy một bức tranh rất rõ về hướng đi của AI hiện tại: AI đang rời khỏi các benchmark tĩnh để tiến vào thế giới hành động thực tế — robot phải chạy nhanh hơn, agent phải tìm tài liệu tốt hơn, coding assistant phải hiểu cả repository, còn mô hình huấn luyện phải học được từ phản hồi tinh vi hơn là chỉ đúng/sai. Dưới đây là phần phân tích 10 paper nổi bật, tập trung vào 4 câu hỏi cho mỗi bài: bài toán là gì, ý tưởng chính, điểm mới, và ứng dụng thực tế . 1) HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone Bài toán: Trong robot manipulation, dữ liệu demo từ con người thường dễ thu thập nhưng chất lượng không đủ ổn định để triển khai thật. Nhiều hệ thống vẫn phải dựa vào dữ liệu bổ sung, tinh chỉnh trên robot, hoặc pipeline phức tạp mới đủ dùng ngoài đời. Ý tưởng: HiFi-UMI hướng tới việc học policy thao tác chỉ từ dữ liệu UMI độ trung thực cao . Tức là thay vì bù đắp bằng nhiều nguồn dữ liệu hỗn hợp, tác giả tập trung nâng chất lượng dữ liệu gốc và thiết kế cách học để policy có thể triển khai trực tiếp. Điểm mới: Điểm đáng chú ý là triết lý “ high-fidelity data alone ”. Đây là một phản đề thú vị với xu hướng “càng nhiều dữ liệu càng tốt”. Bài báo ngụ ý rằng với dữ liệu đủ chuẩn, ta có thể giảm đáng kể phụ thuộc vào fine-tuning tốn kém hoặc domain adaptation phức tạp. Ứng dụng thực tế: Các tác vụ như gắp đặt vật thể, lắp ráp đơn giản, thao tác trong môi trường gia dụng hoặc kho vận. Nếu cách tiếp cận này thực sự bền vững, nó có thể giúp doanh nghiệp triển khai robot nhanh hơn vì giảm chi phí thu thập và hợp nhất dữ liệu đa nguồn. 2) TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM Bài toán: Vision-Language-Action (VLA) rất hứa hẹn cho robot, nhưng thường quá nặng để chạy real-time. Muốn robot phản ứng mượt,

2026-07-30 原文 →
AI 资讯

16 Redesigning my Portfolio Website

Published on Aug 18, 2025 A New Era of AI-Powered Coding Begins I have installed Cursor on my laptop this weekend, and I am amazed at how much it speeds up my coding. I have a new debugging buddy!! This week, I have made several updates to the Portfolio website. The Challenge: When OpenAI Falls Short In my previous post, I shared the excitement of implementing a chatbot based on ChatGPT for my portfolio website. The initial experience was promising - I successfully created content embeddings and integrated them with OpenAI's API. However, as many developers know, relying on a single service provider can lead to unexpected roadblocks. When my OpenAI account encountered issues, I faced a critical decision: abandon the chat functionality or find an alternative solution. I chose the latter, embarking on a journey that would transform my portfolio's AI capabilities and teach me valuable lessons about building robust, fallback-ready systems. The Migration: Embracing Open Source AI The transition from OpenAI to Hugging Face wasn't just a simple API swap - it was a complete architectural evolution. Here's what I learned: 1. Model Selection Complexity Finding the right model on Hugging Face proved more challenging than expected. After testing several options: microsoft/DialoGPT-medium - No inference provider available gpt2 and distilgpt2 - Limited conversational capabilities Qwen/Qwen3-4B - Perfect fit with the nebius provider 2. Database Architecture Evolution The migration also prompted a database upgrade from MongoDB to Neon PostgreSQL. This wasn't just about changing providers - it was about building a more scalable, production-ready foundation for my portfolio. Technical Implementation: Building Resilience Streaming Responses for Better UX One of the most significant improvements was implementing streaming text responses. Instead of waiting for complete AI responses, users now see text appear word-by-word, creating a ChatGPT-like experience: // Streaming implementation

2026-07-28 原文 →
AI 资讯

Top AI Papers on Hugging Face - 2026-07-25

10 paper AI nổi bật nhất trên Hugging Face hôm nay: từ agent tự cải tiến đến benchmark cho “active observers” Hôm nay mình tổng hợp 10 paper đang được upvote cao nhất trên Hugging Face. Danh sách này khá thú vị vì trải rộng nhiều hướng rất “nóng”: deep research agent, hậu huấn luyện mô hình lớn, embodied visual tracking, knowledge graph cho giáo dục, self-distillation cho vision, diffusion language model, đánh giá spatial cognition, sinh video dài, retrieval vượt khỏi “relevance”, và benchmark cho tác tử quan sát chủ động. Bài viết này không đi quá sâu vào chi tiết toán học, mà tập trung trả lời 4 câu hỏi cho mỗi paper: Bài toán là gì? Ý tưởng chính là gì? Điểm mới nằm ở đâu? Ứng dụng thực tế ra sao? 1) AREX: Towards a Recursively Self-Improving Agent for Deep Research Paper : 2607.21461 GitHub : https://github.com/VectorSpaceLab/arex-model Project : https://vectorspacelab.github.io/arex-model/ Bài toán Các “deep research agent” hiện nay có thể tìm kiếm, đọc tài liệu, tóm tắt và lập báo cáo, nhưng vẫn có một giới hạn lớn: chúng chưa thực sự tự cải tiến theo vòng lặp . Phần lớn agent chỉ chạy theo pipeline cố định hoặc được tối ưu thủ công. Ý tưởng AREX hướng đến một agent có khả năng đệ quy tự cải tiến . Nghĩa là agent không chỉ làm nghiên cứu, mà còn biết đánh giá kết quả của chính mình, tìm điểm yếu, sửa chiến lược, rồi chạy vòng tiếp theo . Ta có thể hình dung AREX như một “nhà nghiên cứu AI” gồm nhiều vòng: lập kế hoạch nghiên cứu, truy xuất thông tin, tổng hợp, tự phản biện, tinh chỉnh chiến lược cho lượt sau. Điểm mới Điểm mới quan trọng nằm ở từ khóa recursively self-improving . Nhiều hệ agent hiện tại có “reflection”, nhưng reflection thường chỉ là một bước phụ. AREX có vẻ đẩy ý tưởng này thành trung tâm kiến trúc , biến cải tiến lặp thành cơ chế vận hành chính. Nếu làm tốt, đây là bước tiến từ “agent biết dùng công cụ” sang “agent biết cải thiện cách dùng công cụ”. Ứng dụng thực tế Trợ lý nghiên cứu khoa học Phân tích thị trường, pháp lý, tài chính Tự động

2026-07-25 原文 →
AI 资讯

Building Perri: A Comic Strip Generator

Meet Perri Comic Generator , a lightweight, single-panel comic creator that merges LLM-driven storytelling with real-time diffusion models. By pairing an Gradio frontend with a high-performance backend, Perri orchestrates a seamless pipeline: it takes a simple story seed, structures it into a panel description, generates the art, and burns the dialogue right onto the final image. The best part? It achieves all of this without massive, resource-heavy infrastructure. Every AI model under Perri's hood is under 32 billion parameters , proving that you don't need giant, compute-heavy models to build something amazing. Here is a look inside the architecture and tech stack that powers Perri. The Technical Architecture Perri is built using a clean separation of concerns, splitting the heavy lifting of generation away from the user interface. 1. The Frontend ( app.py ) Built using Gradio 6.16.0 , the frontend provides a sleek, user-friendly interface for inputting story seeds. To match the creative spirit of comics, the UI utilizes a custom theme, incorporating a vintage aesthetic complete with star-twinkle CSS overlays. The frontend's main jobs are: Capturing the user's initial prompt. Shipping the payload to the backend infrastructure via secure API requests. Decoding the backend's response—a Base64-encoded JPEG—and rendering it within the Gradio image component. 2. The Backend Orchestrator ( orchestrator.py ) The orchestrator acts as the brain of the operation, executing three distinct phases in the lifecycle of a single comic panel: Script Generation: It refines the user's raw prompt into a highly structured visual script and dialogue snippet using meta-llama/Meta-Llama-3-8B-Instruct . Image Generation: It passes the visual description to stabilityai/sdxl-turbo to synthesize the retro comic art. Dialogue Overlay Composition: Instead of relying on separate text captions, the orchestrator dynamically draws the generated dialogue directly onto the JPEG image, ensuring an au

2026-06-16 原文 →