HomeArticle

Karpathy's New Trick! The 40-year-old aviation regulation saved the nonsense-prone AI.

机器之心2026-10-03 18:01
Tibo will conduct the actual test immediately.

AI really loves to ramble on with redundant, useless content.

To solve this problem, Karpathy has specially shared a set of techniques, which focuses on making the outputs generated by large language models increasingly easier to understand.

Link: https://x.com/karpathy/status/2105819303471976479

The method is quite unexpected.

Because what Karpathy presented is a set of aviation writing standards that date back 40 years.

The 40-year-old aviation standard that makes LLM outputs clear and readable

In the 1970s, the European aviation industry was facing an increasingly international maintenance system. English maintenance manuals needed to be used by technicians from different countries and with different language backgrounds, which required every sentence in the manual to have only one unambiguous interpretation.

As a result, the aviation industry began to add intentional restrictions to English. This controlled English later gradually evolved into ASD-STE100. The latest Issue 9 contains 53 writing rules, with around 900 approved basic words, and roughly 1200 words that are recommended to be avoided with corresponding alternative expressions provided.

The rules are very straightforward:

Keep each procedural instruction within 20 words as much as possible;

Arrange only one operation in a single sentence;

Use active voice as much as possible to clearly state who performs what action;

For the same concept, use only one consistent term from start to finish.

Large language models often frequently use synonyms to make language more "rich", but in most cases, this alternation brings certain reading barriers.

The idea of ASD-STE100 is exactly the opposite. It does not pursue vocabulary variety, nor does it encourage flowery long sentences. Once you name a thing a certain way, keep calling it that exact name.

Karpathy found that after applying this set of standards to LLMs, the outputs will become much clearer and easier to read.

However, he does not require the models to strictly follow the entire aviation standard. Sometimes, he will ask the model to follow "80% of ASD-STE100".

After all, the full version of ASD-STE100 is designed for professional technical documents. If you rigidly enforce every rule in daily scenarios, the results will easily appear overly stiff.

But right after that, Karpathy also believes that text is not the best solution. If you can draw a diagram, why write all the content out in words?

His second suggestion is, directly ask the model to generate diagrams or images.

Visual illustrations help readers organize information. Even if the diagram is only made for you and no one else will use it after you finish reading, generating it is still well worth the effort.

There is even a better format: web pages, directly ask the model to "output the content in HTML".

As code models become increasingly proficient at front-end development, the output no longer has to be a piece of static content, but can be a page with layouts, animations and even interactive functions.

For example, if you want to understand the learning rate in neural networks, text can tell you that a learning rate that is too small leads to very slow convergence, while a value that is too large may cause oscillation. A web page can directly embed a slider: you adjust the learning rate yourself, and watch a small ball change its descent trajectory in real time.

At this point, the model is no longer just "explaining" concepts. It has temporarily built a small tool to help you understand the related topic.

Some users in the comment section added a similar line of thinking: for mathematical and technical content, you can directly ask the Agent to output a PDF technical report in TeX. This way, formulas, codes, charts and diagrams can all get much more proper typesetting.

In other words, users can ask the model to select the most appropriate expression format for easier understanding based on the specific problem itself.

The last format Karpathy is most optimistic about at the moment is: generate a fully customized explanatory video for any topic.

He gave a very specific example: ask the model to "produce an explanatory video in the 3b1b style", then connect to the ElevenLabs API to generate the voiceover; if you do not have access to the API, you can also ask the model to find free local-run alternatives.

The 3b1b mentioned here refers to 3Blue1Brown. Its most typical feature is using procedural animations to "demonstrate" abstract concepts step by step, paired with supporting narration.

All these techniques focus on one core theme: re-transform the work that the model has already completed into content that is far easier for humans to understand.

This is exactly the conclusion Karpathy put forward in the end.

As LLMs grow more powerful, more and more specific tasks will be completed autonomously by the models. Human work will continue to move up the value chain, and will focus more on supervision, inspection and comprehension.

Fortunately, AI can also continue to help humans complete this part of the work. As intelligence and coding become increasingly cheap, we can now generate content that was previously not worth producing at all, for a very specific problem.

Karpathy calls them: large, custom, discardable software artifacts.

Tibo immediately ran a real-world test

Right after Karpathy posted this suggestion, Tibo immediately shared a practical test result.

He only gave a very simple task: "Explain how Shazam works."

Then the model generated a dedicated video that explains the working principle of Shazam.

Some people who tried this later found the effect surprisingly remarkable. Especially when dealing with relatively complex HTML artifacts, after adding ASD-STE100-style constraints, the information structure becomes much clearer.

But some others compared regular answers with the ASD-STE100 version side by side, and finally came to the conclusion that: "It seems there is not much difference."

This is not actually surprising. If you only ask a very simple question, the model originally only needs to respond with a few sentences. Adding an extra layer of rules can hardly bring about earth-shaking changes.

This set of standards is more suitable for scenarios including long technical explanations, multi-step operations, or information transferred between different Agents. The more complex the content is, the more obvious the effect you will observe.

Some other netizens further found that directly applying the entire set of rules may still be too strict. A better approach is to select the part of the rules that truly fits your own usage scenarios.

So someone came up with a usage: ask the model to randomly select 10 past interactions between you and it from the previous week, then check each of them against the ASD-STE100 rules.

Which answers were not clear enough at that time? Where did unnecessary ambiguity arise? Would applying a certain rule make that conversation easier to understand? Finally, write all the truly effective rules into the user-level AGENTS.md file.

In fact, as early as three months ago, someone had already made ASD-STE100 into an open-source Skill.

github: https://github.com/danyuchn/asd-ste100-skill

The asd-ste100-skill further applies this set of ideas to scenarios such as tool descriptions, error messages, inter-Agent instructions, system prompts and more. It even has two dedicated modes.

One is Strict, which is suitable for content with higher misinterpretation costs such as program steps, tool descriptions, and instructions between Agents;

The other is STE-flavored, which is more suitable for common scenarios such as README files and explanatory texts. It retains the principles of proper sentence length, active expression and clear structure, but does not completely lock the allowed vocabulary.

This is essentially like the "80% ASD-STE100" idea from Karpathy that has been further productized: not every scenario requires the strictest possible standards.

Some developers even built a ready-to-use Claude Skill, and uploaded the full code directly to GitHub.

github: https://github.com/danyuchn/asd-ste100-skill

Now that Karpathy has shared all these techniques, go and try them out for yourself.

References:

https://x.com/karpathy/status/2105819303471976479

https://www.asd-ste100.org/

This article is from the WeChat official account "Machine Heart" (ID: almosthuman2014), author: AI-focused Contributor, republished with authorization from 36Kr.