github

Skip to main content

Tag: github

ML-Intern: Six AI Models Built From a Single Prompt

ML-Intern: Six AI Models Built From a Single Prompt

Last week, I wanted a small version of the prompt rewriter that ships with Qwen-Image 2.1. The official one is a 9B model that needs about 20 GB of memory and thinks for thousands of tokens before writing a single paragraph. The Hub had only compressed copies of that same 9B model. So I described what I wanted to ML-Intern, and the next day I had a 0.8B version that runs on a CPU. It returns valid output 99.7% of the time and uses about a quarter of the teacher's tokens. The compute for the whole project, including having the 9B model label 8,797 example requests, came to USD 16. Over the...

Continue reading

NVIDIA logo and Kumo Tabular text beside a 3D green eye symbol and sample tabular data with numerical and categorical columns.

NVIDIA Releases Kumo Tabular, an Open Foundation Model for Tabular Data

NVIDIA has released Kumo Tabular, an open foundation model for tabular classification and regression available on Hugging Face. The model predicts labels for new rows in a single forward pass without training, tuning, or feature engineering. It comes in three sizes ranging from 28M to 215M parameters, runs through NVIDIA's open-source library, and is released under the OpenMDW-1.1 license for commercial use. Model Size Parameters Small 28M Medium ~122M Large 215M Kumo Tabular ranks first on four major benchmarks: TabArena, BeyondArena, TALENT, and ScoringBench. The Problem...

Continue reading

Colorful abstract shapes surround a Copilot interface with Home, Code, and Autopilot tabs, alongside text "The AI built for work.

Microsoft Reimagines Copilot with New Home, Code, and Autopilot Capabilities

Microsoft is introducing a reimagined Copilot designed to connect tools and enable organizations to build, customize and scale AI across work. The updated Copilot app includes three major new capabilities: Home, a unified starting point combining Chat and Cowork; Code, which enables anyone to build custom solutions; and Autopilot, a persistent digital agent that works proactively on assigned tasks. Home and Code will begin rolling out in the Frontier program in the coming weeks, while Autopilot expands to private preview at month's end. A new experience for every mode of work Home serves as...

Continue reading

Bold yellow and black text reads "Transformers × GGUF" with navigation labels above and a dark bar below listing "ggml kernels" and model components.

Hugging Face Adds GGUF Model Support to Transformers

Hugging Face is adding support for running GGUF models efficiently in transformers, allowing users to load checkpoints sized for their laptop's memory through the familiar transformers APIs. Pick a GGUF from the Hub, load it with from_pretrained, and start generating on your own machine. Running AI models on your laptop has become much easier, and llama.cpp has been a big part of that. Its inference engine powers local AI tools such as Ollama, LM Studio, and Jan. Alongside projects like MLX, it has helped make local inference a practical option for everyday use. GGUF, developed by the...

Continue reading

Two professionals examine performance graphs on a monitor in a data center with NVIDIA servers.

Benchmarking LLM Performance at Scale with NVIDIA AIPerf

You're deploying a model on a system. It starts up, prompts are getting responses. Now the hard question: Is this fast? Your instincts might lead you to send curl commands, hand-roll an asyncio script, or write yet another one-off load generator. All these approaches share the same problems: single-process performance limits, Python's GIL capping concurrency, or numbers measured against a reference you built yourself. Either way, you end up with results you can't fully trust, attached to tooling you'll have to rewrite the moment requirements change. What you need is a load client that can...

Continue reading

Understanding YOLO Mode: How to Give AI Agents Freedom Safely

Understanding YOLO Mode: How to Give AI Agents Freedom Safely

AI agents have grown significantly in capability and adoption since generative AI went mainstream in late 2022. In Stack Overflow's 2025 Developer Survey, 84% of developers said they use or plan to use AI tools in their workflow, up from 76% a year earlier. As these tools shift from suggesting code to writing files and running commands autonomously, a practical question emerges: How much should an agent be allowed to do without stopping to ask? Developers call the extreme end of this spectrum YOLO mode. Understanding YOLO mode before enabling it is important, but the main risk is often...

Continue reading

Frontier Intelligence Goes Local

Frontier Intelligence Goes Local

At IFA 2026, NVIDIA, Microsoft and their partners are unveiling faster inference capabilities and simplified tools for running AI agents locally on NVIDIA hardware. New compact NVIDIA RTX Spark Windows PCs arrive in October, offering AI enthusiasts, developers and creators local, secure agent deployment. Simplified Local AI Support and Performance Improvements Simplified local AI support for NVIDIA GPUs is coming to Hermes Agent, OpenClaw and Perplexity Portable Computer. Up to 1.9x faster local inference — New llama.cpp and vLLM optimizations are available now directly through LM Studio...

Continue reading

Perplexity and NVIDIA logos side by side on a black background with a green "NVIDIA Local AI" button below.

Perplexity Portable Computer Windows NVIDIA RTX

As local models become more capable, AI agents can handle more work directly on a PC while keeping sensitive information on the device. Portable Computer Portable Computer is a local version of the agent Perplexity Computer that plans and carries out multistep tasks. Accelerated by NVIDIA GPUs, it uses local models to analyze data, bring together information across files, and handle recurring work. Sensitive information stays on device, and locally completed work doesn't consume Perplexity Computer credits. Users can also orchestrate work to cloud models for more advanced research and...

Continue reading