HomeArticle

Are large models getting "smaller"?

万点研究2026-07-22 09:41
All large model manufacturers are focusing on lightweight deployment and launching scaled-down mini versions. What is the reason behind this?

WAIC2026 has just concluded, and the large language models from tech vendors seem to have taken an unexpected turn.

At last year's conference, enterprises were still competing over model parameters and benchmark scores. This year, the focus is no longer on "how big the model is," but rather "what the model can do and whether it can generate revenue." A closer look reveals a surprising trend: all manufacturers are pushing for lightweight deployment and launching scaled-down mini versions. What is driving this shift?

Think back: how long has it been since you stared in frustration at an unresponsive voice assistant in an elevator with poor cellular signal? Or maybe just yesterday, you drove through a tunnel and your navigation suddenly went silent. Or on a flight, you tried to organize your photo album, only to find the "one-tap object removal" function grayed out due to no internet connection. We seem to have grown accustomed to the idea that true intelligence must always rely on that distant "cloud."

But now, a quiet revolution is underway.

Those once "super brains" — large language models and multimodal models that used to require an entire building of hardware and consume as much power as a hydroelectric station just to barely run — are now being ingeniously "squeezed" into smartphones, cars, personal computers, and even a humble robotic arm in a factory. Researchers and engineers are sparing no effort to free intelligence from network cables and signal towers, making it wake up instantly, anywhere, anytime, right in the palm of your hand with zero latency. This is like fitting a "mammoth" with hundreds of billions of parameters into a palm-sized "refrigerator," while ensuring it stays alive, capable of thinking, and performs its tasks more precisely and efficiently than in the wild.

This is not science fiction — it's a real "tech spectacle" unfolding right now. What drives this is not the arbitrary decisions of a few individuals, but an industry consensus forged from the genuine demands of the market.

The Cloud-Bound Super Brain Is Descending to the Physical World

At the end of 2022, ChatGPT burst onto the scene, astonishing the world with poems, functional code, and well-written academic papers. Since then, everyone has learned a staggering fact: it relies on a computing cluster of tens of thousands of top-tier NVIDIA GPUs, with massive computational power supporting every conversation we have. This model worked perfectly for a long time — until the real world delivered three harsh wake-up calls to technologists.

Each wake-up call was more painful, and closer to the truth, than the last.

The first wake-up call came from smartphones.

Remember the frenzy in the smartphone industry from late 2023 to early 2024? At the Samsung Galaxy S24 series launch event, the spotlight no longer fell on camera megapixels, but on the AI feature called "Circle to Search." Draw a circle freely on the screen around a bag a blogger is carrying, a landmark in a photo, or a key point in a text, and your phone will immediately tell you what it is, where to buy it, and how to get there — no browser redirection or text input required. Behind these features, Qualcomm Snapdragon and MediaTek Dimensity chips were heavily promoting their ability to run 10-billion-parameter large models locally.

What's so special about running large models locally? What does it mean for ordinary users? Is it just another unremarkable marketing gimmick?

The video of your child's first steps you shot last night, the screenshots of your stock account and bank statements, and your private, intimate chat records with your partner — none of this data should ever be uploaded to a single byte in the cloud. The smartphone has almost become our most intimate digital organ, and it is the most relentless, uncompromising force pushing large models to "slim down."

Then came the automotive industry — a wake-up call rushing in at full speed.

Imagine this scenario: you're driving an electric vehicle equipped with the latest smart cockpit system at 120 km/h on the highway. You casually say, "I'm a bit sleepy, give me something refreshing, but don't blow air directly at my face." How many layers of meaning does this sentence contain? First, it needs to recognize that you are "sleepy," then understand your need to "feel refreshed," which likely requires coordinating the air conditioning, fragrance system, and music. Finally, the instruction "don't blow at my face" demands precise zone control of the air vents, requiring a near-intuitive understanding of the physical world. This entire sequence of actions requires the cockpit system to understand and execute your request accurately through a natural, conversational interaction — rather than rigidly responding with "Hi XX, turning on AC and music..."

From a technical perspective, if this sentence were uploaded to the cloud as-is, processed by a hundred-billion-parameter large model, and then sent back to the car as commands, the round-trip latency could exceed 0.5 seconds. In an office, 0.5 seconds is nothing, but at 120 km/h, you could travel nearly 15 meters in that time. When end-to-end AI large models start handling real-time autonomous driving decisions, every millisecond matters for safety — latency and reliability are governed by the laws of physics, and no matter how powerful the cloud is, it cannot overcome the speed of light or unstable cellular signals.

Cars must be mobile intelligent agents on wheels that can think independently. For safety reasons, to think and act faster and smarter, their "brain" has to reside inside the vehicle itself.

The third wake-up call came from factories.

If you ever get the chance to visit the so-called "lights-out factories" in the Yangtze River Delta and Pearl River Delta, your entire perception of traditional manufacturing will be completely overturned. Take this scene: high-speed cameras scan every electronic component on the assembly line at hundreds of frames per second. In the past, this massive amount of data had to be transmitted via industrial Ethernet to a central server for analysis — not only causing latency, but also meaning that any network jitter or server overload would bring the entire production line to a halt, waiting for the "brain" to send back instructions. Every second of downtime translates to real financial losses.

Now, in edge computing modules and various industrial PCs embedded with AI chips, lightweight visual inspection large models are running directly on-site. They are offline, do not rely on any external signals, and only focus on the products on the assembly line in front of them — making millisecond-level pass/fail judgments with speed and precision far beyond human capability: qualified, qualified, qualified... Reject any part with even a micron-sized scratch! This is like giving every machine the independent, focused, and tireless soul of a veteran quality inspector. The industrial intelligent brain is rapidly shifting from central server rooms to every robotic arm, every drill bit, and every sensor on the factory floor.

These three wake-up calls, from the dimensions of privacy security, physical limits, and cost efficiency, lead to an unambiguous conclusion: True intelligence cannot exist only as a distant cloud. It must flow like oxygen in the bloodstream within every end device, standing by at all times.

This directly spawns a seemingly unreasonable demand: can engineers perform an unprecedented "weight-loss surgery" on the super large model — which currently runs on hundreds of H100 GPUs in the cloud and consumes as much power as a small town — to create a scaled-down mini version, and then elegantly, stably, and powerfully fit it into a smartphone, a car, or a palm-sized edge computing box?

Market demands are already knocking on the door. Inside, a global competition to "slim down" large models has officially kicked off.

The Global "Weight Loss" Contest: How to "Liposuction" a Large Model

Slimming down a model cannot be done through crude, reckless surgery — that would only leave you with a paralyzed, non-functional system. Model developers have quickly reached a consensus in their exploration: what we want is not a emaciated, fragile invalid that collapses at the slightest breeze, but a highly concentrated, energetic, sharply defined "special forces soldier."

Exactly — we need an expert deployed to perform specific tasks. That is the direction. In this smoke-free "tech spectacle," different players are taking vastly different approaches.

Let's first look at the U.S. West Coast, where old rivals OpenAI and Google have each developed their own distinct paths.

OpenAI's strategy is the most straightforward and commercial: vertical grading, tiered pricing based on capability. In the cloud, one of the most powerful AI brains, GPT-4, guards the high ground. This colossal model explores the limits of intelligence, working on mathematical conjectures and writing sonnets. But OpenAI soon realized that among billions of global users, the most common daily queries are actually things like "help me name a fried chicken restaurant," "summarize this meeting note into three points," and "how do I fix this buggy line of code." Having a highly paid, busy doctoral supervisor answer elementary school students' arithmetic problems is not only a waste of intellectual resources, but also a disaster for the business model.

Then came GPT-4o mini. Its official announcement made no secret of its positioning: better cost-performance, lighter, and faster. It delivers extremely fast responses, with call costs plummeting by over 60%. While it may not be as knowledgeable as its big brother GPT-4, it performs surprisingly efficiently on high-frequency, massive tasks such as daily conversations, text summarization, customer service Q&A, and code assistance — the difference is barely noticeable. It's like having a Nobel Prize-winning mentor and a brilliant PhD senior fellow apprentice: the apprentice quickly and clearly solves all your daily questions and homework problems, available whenever you need them, while the mentor focuses on critical thesis defenses and strategic direction setting a few times a year.

This is stratification, this is efficiency.

Google, meanwhile, follows the path of "full-size coverage, hardware-software integration," full of the characteristic romance of a "hardcore engineer." From its initial release, their Gemini model was clearly divided into three versions: Ultra, Pro, and Nano. Its ambition is obvious: to be as flexible as Sun Wukong's golden cudgel, able to grow large or shrink small. Ultra pushes the boundaries of human knowledge and dominates academic benchmarks; the Pro is the capable "middle class," integrated into the full Google Workspace suite to help you write emails and create presentations. The one that sends chills down competitors' spines, however, is the most unassuming version: Gemini Nano.

Gemini Nano was specifically designed, even custom-tailored, to run natively on Android mobile devices such as smartphones and tablets. Google's engineers used state-of-the-art distillation algorithms and quantization techniques to compress core, high-frequency AI capabilities — like fraud SMS detection, smart reply suggestions, and offline speech transcription — into a tiny, highly efficient offline engine. The result? When you use a Pixel 8 Pro powered by the Tensor G3 chip, even in airplane mode, it can still accurately filter spam messages and instantly generate a smart text summary of your voice memo. All computations happen right on the chip in your hand, with not a single bit of your data ever leaving the device.

But these tech giants are not invincible.

Back in China, innovation is surging, spawning more diverse and down-to-earth solutions.

When it comes to miniaturizing large models and making them fully open-source, DeepSeek is impossible to ignore. Few people know that the founder of this company has a background in quantitative trading. Perhaps this is what gives their corporate DNA an almost obsessive pursuit of "efficiency" and "cost-effectiveness." DeepSeek's strategy is more extreme: not only did they release top-tier cloud large models rivaling GPT-4, but they also vigorously launched a matrix of delightfully tiny "micro" models at 1.3B and 6.7B parameters.

These "little guys" specialize in code generation and mathematical reasoning. How impressive are their results? A 6.7B small code model can solve specific Python function writing and algorithm competition problems at blazing speeds, with performance comparable to models dozens of times larger. Most crucially, they are completely free, open-source, and allow commercial use. This means a three-person geek team working out of a rental apartment in Huilongguan, Beijing, can run an AI colleague on a self-assembled server worth a few thousand yuan — an AI with capabilities on par with a three-year experienced algorithm engineer, tirelessly helping them write code and debug 24/7. Deepseek has turned the grand slogan of "AI democratization" directly into a tangible productivity tool for small teams — that's the real game-changing power.

If DeepSeek's philosophy is about "becoming smaller," then another rising star, Kimi, has redefined what it means to be "lightweight."

In fact, Kimi's founder Yang Zhilin and his team have been asking a fundamental question from day one: what is the core bottleneck restricting current AI applications? Their answer is "memory" — or more precisely, "working memory." Human thinking is linear and continuous. After a two-hour brainstorming session, the final summary requires recalling information from every previous sentence. That is the invaluable value of long context windows.

Earlier AI models mostly had "goldfish memory": once a conversation got long, they would start rambling nonsense and forget the initial role setting of "you are my legal counsel." Instead of rushing to launch a 7B end-side small model like everyone else, Kimi focused all its technical resources like a laser on breaking the physical limits of context windows.

While other models were still showing off their ability to process dozens of pages of PDFs, Kimi could ingest a 3-hour audio recording, a 100-page financial report, and 50 academic papers in one go — then accurately answer questions like: "According to the previous board resolution and the latest competitor analysis report, there are three inconsistencies in our current strategy..."

This is Kimi's unique philosophical solution to "becoming lightweight."

What other models need a complex system of "large model + vector database + retrieval-augmented generation tool + multi-agent" to accomplish — the "read-understand-summarize" long text task — Kimi can handle all by itself. This is still "lightweighting," but it does not reduce the model's size — it reduces the complexity of the entire application architecture and development costs. This uses algorithm-level depth to replace the breadth of stacked system components, an Eastern wisdom of "using four ounces to move a thousand pounds" that bypasses the detours imposed by hardware limitations.

And those closest to the hardware ecosystem are the ones bringing all this technology to ordinary consumers.

New players like Facewall Intelligence represent another extreme romantic vision — taking "smallness" itself as a core belief. Their "Little Steel Cannon" series of end-side models, even with only a few hundred million parameters, have demonstrated astonishing capabilities on basic natural language processing tasks. It reveals a future possibility: the AI around us does not all need to be Einstein-level geniuses. A washing machine that knows all garment classification rules and can wash silk shirts with the gentlest, lowest-wear program, or a rice cooker that can accurately identify rice types and freshness and automatically adjust heating power — these are the real "Doraemon" helpers that millions of households truly need. This is the value of specialized experts, and the true market depth of "miniaturization."

You see, whether it's hierarchical stratification and open-source ecosystems overseas, or extreme cost-effectiveness and algorithm innovation in China, all the world's sharpest minds are answering the same question in their own way: how to free powerful intelligence from the constraints of the cloud, and turn it into a truly portable, private, real-time, and tangible capability. All paths eventually converge to three core principles: letting top geniuses act as private tutors, walking the fine line between precision and efficiency, and building models specifically tailored for specific scenarios.

At this point, the industry's future seems bright. But a fundamental question, lingering in every thinker's mind like an unshakable dark cloud, remains unanswered.

Will Smaller Mean Dumber? Unraveling the Core Concern

When the model's size is "shrunk," can it still remain as "smart and sharp" as before?

This concern is completely natural and intuitive, as we all believe in "more power, more results." The more parameters, the denser the neural connections, the more "knowledgeable" and "intelligent" the model — this has long been seen as the iron law of AI development. From GPT-1 with hundreds of millions of parameters, to GPT-3 at the hundred-billion level, to GPT-4 which is rumored to have broken the trillion-parameter mark, the emergence of intelligence seems to be achieved through brute-force stacking of computing power. Now you're telling me that we can cut 99% of the "fat" from a hundred-billion-parameter model, compress it to a fraction of its original size, and still retain its amazing cleverness? That seems to violate the law of conservation of energy, doesn't it?

To this sharp question, the industry's answer is surprisingly candid, yet logically reassuring: if you forcefully compress a model, it will definitely malfunction — but what if we don't train it to be a "doctoral supervisor" from the start, but instead carefully nurture it to become a "specialized expert"?

This is the most fundamental, revolutionary shift in mindset, overturning our understanding of AI development from the past few years. Previously, we kept creating "universal geniuses," but when that genius is deployed to the "one-tap photo retouch" function in your phone's album, 99.99% of its knowledge becomes expensive, memory-hogging "useless baggage." Its only job is to look at a photo and make precise millisecond judgments: which areas are flawed shots? Can we fix a person's closed eyes in the photo? Can we seamlessly erase that random trash can in the background? It doesn't need to understand quantum physics