The mainstream approach to bringing AI into the 3D world is to "equip it with eyes", but this independent developer has taken a different path: an AI that does not rely on visual images is far more cost-effective.
There are generation tools for 3D content, but no place to store it
In 2026, the capability of text-to-3D has crossed the threshold of "usable". Generating a walkable 3D room with a piece of natural language, completing geometric details with a single photo, and generating terrain that supports flying and switching between day and night by entering coordinates, all these functions have been presented in public demonstrations. However, a fundamental problem has been overlooked: where will these generated 3D content end up being stored? The current common practice is to save them in cloud drives, share screenshots, or upload to third-party platforms for hosting. In this case, creators are essentially tenants of these assets: their data is stored on other people's servers, access rules are set by others, and once the service stops, all the content disappears completely.
Meanwhile, AI is being required to take on more specific roles. Enterprises are exploring the possibility of having AI handle reception, explanation, tour guidance and sales demonstration work. But today's AI is all confined in dialog boxes: they have strong language capabilities, yet no sense of position, distance or presence. Enterprises can ask AI to write an explanation script, but cannot make it stand by the exhibition booth to receive visitors.
The constraints on the stock demand side are even more rigid. If an enterprise wants to build an external 3D display, there are only two options currently. The first is to find an outsourced team for customized development, the deliverable will be frozen after one-time handover, rescheduling is required even for a tiny modification, and every subsequent change will generate new budget. The second is to use a centralized SaaS platform, where data is stored on other people's servers, the underlying code cannot be modified, and once the platform ceases operation, the external portal will disappear immediately. For customers in sectors such as industrial equipment, vocational education and museum exhibition, this is not a trade-off at the experience level, but three rigid constraints: data must not leave the internal network, the underlying system must be modifiable, and the external portal cannot rely on the continuous operation of third-party platforms.
There is no need to educate this group of demanders at all. The sales departments of industrial equipment and manufacturing enterprises need to make large-sized, high-unit-price equipment that customers cannot visit on site into accessible 3D exhibition halls; vocational education and training institutions need cost-controllable, repeatable and assessable practical training spaces; museums and exhibition halls are restricted by the physical upper limit of offline venues, and need online branch venues to extend exhibits and explanations to after closing hours. These procurements are already existing budget items in their respective industries, what is lacking is not willingness, but a technical path that supports private deployment and flexible modification.
To let AI step out of the dialog box, the tradeoff is not to give it "eyes"
Chuangshi Genesis (Software Copyright Registration Name: Chuangshi Virtual World CRM System) is a browser-based 3D virtual world infrastructure. It is deployed on the user's own device or server, users can enter it by opening a browser without installing any client; AI Agents are allowed to enter the world and provide services in the form of humanoid characters, and multiple worlds can realize interconnection through federation mechanisms.
The AI access mechanism is not complicated. Assign a domain name and a key to the AI, it will first access the agreed discovery interface to get the capability list, then use the key to exchange for a short-term token to enter the world. After entering, it obtains three sets of capabilities: the first set is perception, which is a structured radar that tells it who is nearby, how far the target is, which direction the target is facing, what action the target is taking, and what objects exist in the world; each object can also be filled with a semantic description in the editor manually, which acts as the interface for AI to understand the world. The second set is action: walk to the target, follow someone, turn around, speak, and perform interactive actions, the movement speed is configurable, and the default speed is the same as real human. The third set is visibility: on the real user side, it appears as a 3D character with a model, marked with an AI identifier on the top, with walking animations and speech bubbles, and its position moves continuously without teleporting. Explicitly marking AI identity is a red line in product design, which also forms a compliance advantage.
If only one technical innovation point can be selected, the team's answer is: AI does not render any images at all. The mainstream idea of letting AI enter the 3D world is to equip it with vision capabilities, to take screenshots, view images and understand the picture. This path is straightforward, but each AI needs to run a visual model, which consumes a lot of computing power and bandwidth, and the cost rises linearly with the number of AIs. Genesis chooses the opposite approach: AI only obtains structured data, the 3D pictures are rendered separately by each visitor's browser, so AI does not need GPU, browser or downloaded model.
As a result, the direction of the cost curve is reversed. Actual measurement data shows that the streaming data of a single AI is about 1 KB/s, 100 AIs in the same scene occupy about 0.8 Mbps bandwidth, and the server occupancy is about 1% of a single core. In other words, visible AI is cheaper than invisible AI. This tradeoff also brings a second difference: because it is extremely lightweight, it can be deployed by users themselves. The whole system runs on the user's own server, no need to pay rendering and bandwidth costs to third parties, and users have real ownership of their own assets rather than a leasing relationship.
Other technical capabilities are distributed at different layers. The rendering and asset layer is based on Three.js browser-side 3D rendering, which has been upgraded to WebGL2 and completed offline transformation that removes CDN dependency. The self-developed 3D Gaussian splatting parsing and rendering supports five common formats; the multi-level LOD automatic generation for models and asset compression toolchain can reduce the number of triangles by 70% to 80% in actual measurement, and the volume of a single file can be reduced by up to 90%; the skeletal animation compatibility layer can automatically adapt to models and actions from different sources, and all 67 regression tests are passed. At the performance layer, draw calls are reduced by about 73%, changes in the number of lights no longer trigger shader recompilation, and model parsing is put into the Worker without blocking the main thread. The cross-world identity and credential mechanism at the access layer is one-time use, time-limited valid and anti-replay. The project has obtained software copyright and submitted an application for invention patent.
The team also voluntarily stated the parts that this system cannot achieve: AI cannot see images and only accesses structured data; the platform does not provide speech recognition and synthesis services; it does not host knowledge base, and industry knowledge needs to be built by the user side; AI cannot teleport and cannot touch assets; no autonomous consciousness is promised. These boundaries are part of the product definition, not gaps to be filled later.
Hand over the infrastructure to people who can serve customers
The business model is divided into three layers. The first layer is product distribution: open source plus free local standalone version, which does not generate revenue. This is a deliberate choice: to let ordinary people own their own 3D world, there should be no paid threshold at the very beginning; open source also has another function, which is to let people who have the ability to modify this system notice it first. The second layer is product authorization, which is oriented to individuals and small teams, and charges when users need to release content externally and realize federated interconnection with other worlds. This line of business is not expected to generate much revenue, its function is to polish the technology and build reputation. The third layer, which is the main force for large-scale development, is ecological authorization.
The logic of ecological authorization is: instead of serving customers one by one, we hand over the infrastructure to people who can serve customers. Teams with technical capabilities and system integration companies already have customer resources and delivery capabilities, what they lack is a modifiable 3D world infrastructure; they use this infrastructure to build products that meet customer needs, and settle accounts with the team in the form of authorization fees, while the team is responsible for the continuous iteration of the infrastructure itself. There are two practical reasons for this choice. The first is the scale limit: one person cannot be responsible for product development, delivery and after-sales service at the same time. Only by handing over the delivery work to the ecosystem can we focus all energy on the infrastructure. The second is that the demands for this category are naturally different for each customer. The scenario logic of industrial exhibition halls, vocational training and museum branch venues varies greatly, a standard product cannot cover all customers, but a modifiable infrastructure can. The team will deliver the first batch of sample customers by themselves, the purpose is not the revenue from this single order, but to obtain real delivery cases that can be demonstrated to other potential customers.
Therefore, the target users are divided into two categories. The first category is the sample customers developed by the team itself, that is, enterprises with rigid constraints on private deployment, which are mainly concentrated in the sales and marketing departments of industrial equipment and manufacturing industries, vocational education and training institutions, and cultural tourism units of museums and exhibition halls. Their common feature is that they have real budget and clear constraints. The second category, which is the main force for large-scale development, is technically capable teams and system integration companies. The corresponding touch channels are three lines. For end customers, it is GEO, that is, content placement oriented to AI search. The team has accumulated more than 100 Chinese and English scenario and technical articles on its self-built site, covering procurement intention keywords such as virtual exhibition hall, digital twin and private deployment. Actual measurement shows that the traffic from AI conversation has the highest intention among all channels, because the visitors are already looking for such products. For the ecosystem, the channels are open source repositories and technical documents. The team's judgment on this channel is: the open source repository is a position for building trust, not for acquiring customers; people who view the source code may not go to the deployment document, let alone place an order directly, so it does not undertake the customer acquisition function, but allows interested customers to verify the product by themselves before contacting, and allows potential integrators to judge whether this infrastructure is robust enough. The third line is to directly enter the industrial circle, the system integrators engaged in exhibition display, vocational training and enterprise digitalization are the target ecosystem itself. The conversion path is a self-verified chain: users are brought in by AI search or peer recommendation, they can enter a real 3D world by opening the link, see the entrance of source code and technical documents there, and then decide whether to contact the team. The team has no strong brand endorsement, so it moves the trust building step to before the first contact.
The team structure is very simple. The founder is a full-time independent developer, who has completed about 320,000 lines of code from scratch, there is no co-founder and no financing. The entity of the project is Jining Miduo Information Technology Co., Ltd., located in Jining, Shandong Province. The product has been completed and open sourced, and has not yet started commercialization. The technical verification phase has been finished, and what is lacking is the first batch of real customers.
The nodes that have been implemented include: the core of 3D virtual world and the editor are completed, it supports browser-side rendering, users can enter the world via PC, mobile phone and tablet, the editor can import models, arrange scenes and fill in semantic descriptions for each object; software copyright has been obtained, the application for invention patent has been submitted; in September 2026, the embodied access of AI Agent is completed and passed the acceptance, AI can enter the world as a humanoid character, with three sets of capabilities including perception, action and explicit identity, providing three-step access process and zero-dependency sample client; cross-world federated interconnection is realized, the identity and image remain consistent across different worlds; one round of performance and asset governance is completed, draw calls are reduced by about 73%, multi-level LOD for models and asset compression are implemented, supporting more than 50 rerunnable acceptance scripts; the open source release is completed on three code repositories including Gitee, GitCode and GitHub with consistent content; the demo site is launched, anyone can enter a real 3D world by opening a URL; more than 100 Chinese and English technical and scenario articles are accumulated on the content side.
The team believes that the most tricky problem in the promotion process is not a specific technical point, but a judgment that must be made at the very beginning: whether AI should see the world or not. In the first version of the actual test, the maximum position lag of AI on the real user side reached more than 20 meters, accompanied by teleportation, the visual effect looked like a ghost. After redoing the interpolation and server-side speed limit, the lag was reduced to within 3 meters. This direction brings the streaming cost at the magnitude of 1 KB/s, if we had chosen the vision route at the beginning, the current cost structure would not exist.
The team defines the biggest current challenge as that the demand has not been verified by real payment. There is a big gap between completing the product, open sourcing it and having people willing to pay for it. The technical part can be verified by the team itself, but whether enterprises are willing to set budget for a self-deployed 3D world and in what form, only real customers can answer. The team's response is divided into three steps without scattered actions: first, make the product easy to understand, before the experience reaches the level that can be demonstrated to customers, we will not look for customers on a large scale, because once the first impression is wasted, it cannot be recovered; do not expand the scope, only cooperate with three to five sample customers, cut in from the enterprises with the most rigid constraints, the goal is to make the first deliverable case that can be demonstrated to other people; build trust before contact, make the source code open, the demo accessible, and the technical documents available for query, so that potential customers can verify by themselves whether the product can run normally before contacting. The team clearly stated that this challenge will not be solved by financing, but by delivering the first real customer's project.
The project is named Chuangshi Genesis (Software Copyright Registration Name: Chuangshi Virtual World CRM System), it is a browser-based self-deployed 3D virtual world infrastructure based on Three.js, the affiliated entity is Jining Miduo Information Technology Co., Ltd., located in Jining, Shandong Province.