HomeArticle

Cards are just the starting point, and AI voice recording devices are vying for the personal memory access point: collar, chest or wrist?

雷科技2026-07-29 12:15
The form factor of AI recording devices is nowhere near the stage of convergence.

Now when you buy a personal AI voice recording device, the more troublesome part than choosing a brand is: which type do you want?

Some people stick it magnetically to the back of their phone, some clip it to their collar or wear it on their wrist, while others hope to fit it into earbuds, glasses, pens, or even make it into a ring. In just two or three years, AI voice recording has evolved from a single card form into more than a dozen different shapes. These products vary greatly in appearance, but they run on a similar underlying workflow:

The microphone is responsible for collecting and storing audio, the mobile phone or cloud completes transcription and speaker diarization, then hands the result to the large language model for summarization and to-do item extraction, and finally the output is synced to apps, documents or knowledge bases.

Why is this track so bustling?

For one thing, it is indeed easier to develop than many other AI hardware products. Mature MEMS microphones, low-power audio chips, flash memory and batteries have lowered the hardware threshold sufficiently, and transcription and large model capabilities can also directly call cloud services. A large number of unbranded white-label products from Huaqiang North claim to be "AI hardware", but the AI processing actually mostly happens on mobile phones and servers.

For another, voice is indeed the cheapest and most context-rich data source for AI to understand an individual. We output a huge amount of information every day in meetings, interviews, phone calls and casual chats, including facts, decisions, commitments, emotions and fleeting ideas. Compared with requiring users to constantly type, take photos and organize files, a portable device only needs to listen quietly, and can fill in a large missing part of an individual's life context for AI systems.

So manufacturers have long been selling far more than just "automatic meeting minute generation". From Plaud, DingTalk, Mobvoi to Bee, Omi, Vocci, terms like "second brain", "external memory" and "personal knowledge base" are repeatedly mentioned, hoping to push voice recording devices to become the entry point for personal AI memory.

The problem is that recording a meeting, remembering what happened in a day, and capturing a random thought in passing are all different demands. Even the most popular AI voice recording cards on the market right now mostly solve the demands and pain points of traditional voice recorders and other personal recording devices.

The product form of personal AI voice recording is far from reaching a final convergent state.

Card form wins the first round, but the AI voice recording game has just started

There is no doubt that the most successful product in the current personal AI voice recording market is the AI voice recording card.

Plaud Note basically defines the standard form of this type of product: card-shaped, magnetically attached (usually via a magnetic case) to the back of the mobile phone, records offline meetings with a regular microphone, and uses a vibration sensor to capture mobile phone calls. At the same time, pressing the body can mark key points in the recording, allowing AI to prioritize processing that segment after the meeting.

Plaud Note Pro, Image source: Leikeji

According to the official statement of Plaud, its global users have exceeded 2 million at present, which at least shows that AI voice recording cards have gone beyond the pure conceptual stage.

The advantages of the card form are not mysterious. It is thin enough to be carried together with the mobile phone. Its body is not so small that it cannot fit a battery, flash memory and multiple microphones. When placed in the center of the table during a meeting, it usually captures the voices of all attendees more evenly than a device pinned to one person's chest. When attached to the back of the phone during a call, it can also bypass the iPhone's restrictions on call recording. It covers the three most common recording scenarios: meetings, interviews and phone calls.

This is a very pragmatic compromise.

Precisely because of this, the voice recording card market has quickly become extremely competitive. Mobvoi, the company behind TicNote AI voice recording card, continues to develop its product into a tool for project knowledge base, in-depth research and podcast creation. Notta Memo cuts into the market relying on Notta's existing transcription and international collaborative service capabilities. In addition to 360 AI Voice Recorder, iFlytek has also launched its own AI voice recording card.

Magnetic attachment, dual mode for calls and on-site recording, dozens of hours of recording time, and hundreds of minutes of free transcription quota per month have almost become a standard public template for such products.

DingTalk has taken a more aggressive path. The DingTalk A1 Youth Edition and Flagship Edition not only push the hardware price directly to the price range of mainstream voice recorders, but the A1 Pro launched in April this year is also equipped with a directional microphone and a 2980mAh battery, claiming to support 10-meter pickup range, 180 hours of continuous recording, and even temporary reverse charging for mobile phones.

But what DingTalk really wants is obviously not to sell hardware. After the recorded audio is imported into DingTalk, it can continue to generate meeting minutes, to-do items and analysis reports, and then sync to documents, AI spreadsheets and enterprise knowledge bases. The previously scattered offline sales visits, interviews and customer communications are thus turned into data that DingTalk can continue to process.

Image source: Leikeji

As the voice recording card market becomes increasingly saturated, the focus has gradually shifted from "whether there is AI summarization" to more trivial but practical detailed experience: whether the audio can be clearly captured in noisy environments, whether the system will misidentify the speaker when multiple people speak over each other, whether the original recording can be fully exported, and whether the transcribed text can be integrated into the user's existing workflow.

Hardware differentiation is getting smaller and smaller, and software and ecosystem are more critical.

However, the card form still requires users to remember to bring it, take it out, place it properly, and then press the recording button. For a pre-scheduled meeting, these steps are not troublesome, but when it comes to walking conversations, sudden inspirations, or when the user really wants to record the whole day, it cannot achieve a non-intrusive experience. As a result, we can see more products leaving the mobile phone and being worn on people's bodies.

On collars, chests, wrists, fingers: AI voice recording is still finding its optimal position

No device is more "portable" than a wearable device, and the biggest value proposition of personal AI voice recording is to build a personal AI memory entry based on "portability" and "always on" features, to record our casual remarks, to-do items, sudden inspirations and conversations, and turn a large number of fragmented "contexts" into "personal memories".

So in addition to the voice recording cards that are magnetically attached to mobile phones, more manufacturers hope to integrate AI voice recording functions into head-mounted devices (AI glasses), collars, chests, wrists, fingers... any possible part of the body, or even develop multi-form products.

Many people may not know that in addition to voice recording cards, Plaud has also launched the voice recording pendants NotePin and NotePin S, which can be clipped on clothes, or fitted into lanyards and wristbands, supporting long-term continuous recording. The first-generation NotePin is started by long pressing and vibration feedback, whose feedback is not intuitive. This point was pointed out in the review of the first-generation product by Leikeji:

...The pressing force required is far beyond expectation. The Leikeji team tried many times and found that only hard heavy pressing can trigger the startup, and even so, there is still a certain probability that the device will not respond. It is reasonable to infer that the original intention of the NotePin product team was probably to avoid accidental startup, so they raised the threshold of pressing trigger force. But the final effect is far from satisfactory: users need to press very hard with their thumb to start or stop recording.

NotePin and NotePin S, Image source: Leikeji, Plaud

Interactive feedback is a detail that cannot be ambiguous, so for the new generation product NotePin S, they simply added an ordinary "physical" button, which turns out to be more reliable.

In addition, Anker's soundcore Work can also be clipped on the collar or hung on the chest. Although the single-device continuous recording time is only 8 hours, it can be extended to 32 hours with the charging case. The AI voice recording bean launched by Anker in cooperation with Feishu this year adopts a similar hardware path: the recorded audio will be directly imported into Feishu Magic Note, and then called for voiceprint recognition, real-time minutes generation and knowledge Q&A.

The same type of hardware, when connected to the soundcore App or Feishu, will eventually become two different products. The former is more like an independent AI voice recorder, while the latter is designed to collect offline work information for Feishu from the very beginning. This once again illustrates that the value of personal AI voice recording devices is increasingly dependent on where it sends the voice, or the "context".

Anker AI Voice Recording Bean, Image source: Anker

The advantage of the clip-on device is that it is ready for use at any time. During mobile interviews, exhibitions, store visits and walking conversations, users do not need to take out their phones or find a proper placement position. But the microphone moves with a single person, which is farther from other speakers, and is more likely to pick up clothing friction noise. Wearable pendants may not be more suitable for multi-person meeting scenarios, but they are more "personal" than recording cards.

It is not only through product design that the recording scope is expanded from meetings to the whole day, the more important part is the AI system. For example, Bee Pioneer can be worn on the wrist or clipped on clothes, and is claimed to be usable for 7 days on a single charge. It continuously understands the user's conversations, generates daily reviews, reminders and interpersonal relationship clues. After Amazon acquired Bee, it continued to push the product towards the direction of personal assistant.

For privacy protection, Bee claims that the original audio is only processed in real time and never saved, and only the transcribed text and summary are retained.

This sounds very thoughtful, but it also brings another problem: if the AI mishears the content, users will not even have the chance to listen back to the original audio and check the context. The Verge once tested Bee for several consecutive months, and found that it would confuse personal relationships, and might even mistake content playing on TV for real-life conversations.

Omi has chosen a more open path, opening the hardware, applications and code to developers, allowing users to access local models, personal knowledge bases and automation tools. It is more suitable for people who are willing to build their own "second brain", but its product completeness and usage threshold cannot be compared with an out-of-the-box voice recorder.

Earbuds are also competing for the same position. Whether it is Mobvoi's TicNote Pods 4G or iFlytek's conference earbud series, both have added independent microphones and 4G modules to the charging case, so that recording upload no longer relies on mobile phone networks, and can directly support calls and translation.

Of course, the most eye-catching attempt in recent times is the voice recording ring.

For example, the Vocci Ring designed for long-time recording can start recording with a double tap, and mark key points with a single tap. The Vocci Ring is equipped with a MEMS microphone on its body, claiming to support 5-meter pickup range, up to 8 hours of continuous recording, and the charging case can provide three additional full charges.

Image source: Vocci

The early experience review from TechRadar mentioned that the original audio sounds a bit muffled, but the transcription performance is better than expected. Whether the first-generation product can maintain its performance in real meetings, noisy exhibitions and long-term wearing scenarios still needs to be verified after mass production.

There are also some other "AI voice recording rings", such as Pebble Index 01. After pressing and holding the button, you speak to the ring for a few seconds, the mobile phone completes the transcription locally, and then creates notes, reminders or executes instructions. No subscription is required, and the silver-oxide battery can support several years of fragmented use, but the maximum single recording duration is only 2 minutes, which is only suitable for fragmented daily recording scenarios.

The activation action of the ring is minimal, but the physical limitations are also the most difficult to avoid. It is far from the mouth and the meeting table. If the hand is tucked into the cuff or placed under the table, the captured sound may become muffled; it is very difficult for a single ring to fit a microphone array, large battery and stable connection at the same time. More subtly, the less noticeable the ring is to others, the more users need to seriously handle the issues of recording informed consent and social boundaries.

At least at the present stage, the ring has more potential to become an easy-to-use AI voice shortcut. Whether it can replace recording cards and voice recorders still needs to pass four barriers: audio pickup quality, battery life, user trust and mass production maturity.

Can recording more content really help you remember more?

So, is personal AI voice recording reliable as a "personal AI memory entry"? If we limit its application scope to clear scenarios such as meetings, interviews, courses, sales visits and medical communication, the value of AI voice recording has been fully verified.

For editors of Leikeji, after an interview, being able to quickly search for the original remarks, locate the time point, and then let AI sort out the context can indeed save a lot of time that would otherwise be spent on dragging the progress bar repeatedly. Users can also easily judge whether the product has completed its task: whether the audio is clear, whether the transcription is accurate, whether key viewpoints are missed, and whether the original audio can be played back at any time.

The reason why recording cards are temporarily leading is largely because they first captured these clear, stable and willingness-to-pay demands.

But the problem arises when the concept of "personal memory" is further expanded. Humans rarely remember an event only through sound. Who is speaking, where you are, what you see, and what is displayed on the screen will all change the meaning of a sentence. All-day recording can accumulate a huge amount of text, but it will also include content from TV, passers-by and irrelevant casual chats. The context is not fully completed, but information junk has piled up into a mountain first.

Transcription is only the first layer. Once the speaker recognition function misidentifies the speaker, the large language model may attribute the commitment of person A to person B; if the summary omits a qualifier, the conclusion may be completely reversed. If the system does not save the original audio like Bee does, it