The technical architecture of transform is used to train the language, images, text, limb movements and emotions associated with brain-computer interfaces, which are essentially human consciousness.
I just came back from Shanghai, and the biggest takeaway from this trip is that I got to meet Professor LIU Quanying, who is also working on brain-computer interface foundation models. Her team is currently building a large foundation model dataset for multi-modal data including EEG, MEG and FMRI.
It is similar to MindSoul, but also different. Because we not only have the foundation model, but also the product form of spatial computing MR glasses for end users.
In other words, MindSoul positions its brain-computer interface as a spatial computing accessory for consumers. In 5 years, as the brain-computer interface supply chain matures, we will upgrade it to a final form of hardware with optimized price, thinness and wearing comfort, but this is not achievable at least for now.
A standalone brain-computer interface cannot meet users' needs, no matter for reasons of accuracy, user operation habits or user cognitive barriers.
The realization of "controlling things with mind" requires the joint efforts of brain-computer interface equipment manufacturers, academic institutions and social education, rather than expecting users to master the use of the device immediately after they get it.
For example, for the current motor imagery paradigm that asks users to look at the screen and imagine making a fist with their left or right hand, there is actually a failure rate of more than half. This is not only caused by the electroencephalogram itself, but also by the fact that users do not know how to visualize the action of making a fist, or the specific difference between the fist-clenching movement of the left hand and the right hand.
What MindSoul does is to complete the respective transform training for models covering text, images, speech, emotions, limbs and all ten fingers.
We believe that if consciousness refers to the part of the human brain and independent consciousness that excludes passive and subconscious contents, then this part is exactly what we define as human consciousness.
With the spatial forum feature on MR devices, it can be said that we can predict the future, peek into a person's destiny, and thus predict the trend of the whole world.
The Foundation Large Model of Brain-Computer Interface: MEG and FMRI Remain the Gold Standard Data
At present, MEG devices are still the top choice for non-invasive data collection of brain-computer interfaces. Even with equipment manufacturers of different brands, MEG data still has the highest purity. Meanwhile, test subjects can perform various possible tasks to complete data collection, and the data can be used for training via transformer.
Recently, Facebook open-sourced tribeV2, which is built with FMRI gold standard data. It can collect relevant data of humans completing video, speech and image processing tasks in spatial scenarios to finish transformer training.
Why We Designate Several Specific Tasks Instead of Collecting All Types of Data
Humans have some high-frequency and universal basic capabilities in social life, including text comprehension and expression, image perception, emotional response, memory and thinking, as well as motion control of limbs and ten fingers.
There is no need for our MindSoul team to collect and train all human behaviors, otherwise the task boundary will expand infinitely, and the data scale and development cost will get out of control.
This is similar to the data burial point design of products: product managers usually only set burial points around key behaviors and core events, instead of recording every single action taken by users.
Therefore, more data does not always mean better. What really matters is whether the data can cover key capabilities, have clear labels, and form trainable, verifiable and reusable neural representations. Focusing on core tasks can not only reduce the cost of data collection, labeling and model training, but also make it easier to generate practical research and application value.
Based on this principle, the MindSoul team is currently focusing on five types of active consciousness related capabilities: text, images, emotions, limb movements and fine movements of ten fingers, and is gradually exploring the relationship between these capabilities and memory, intention and thinking activities.
Just like writing this article today, we have started to label the data of the spatial forum on transformer to complete the training of the world model, and at the same time label the data of text, images, emotions, ten fingers, left and right hands, as well as left and right feet on MindSoul.
Spatial Forum World Model + MindSoul Neural Large Model = The Future
For example, the picture below shows the world model labeling of the spatial forum we are working on and the subsequent transformer training process.
The above is some of the forward-looking content I sorted out during my recent stay in Shanghai. The essence of building a brain-computer interface large model is to start predicting the future, because the combination of the world model and the MindSoul neural model is exactly the way to realize future prediction.
That's all for today's sharing.
This article is from the WeChat Official Account "Kevin's Bits That Change the World" (ID: Kevingbsjddd), author: Kevin's Stories, published with authorization from 36Kr.