HomeArticle

The longer the contracts, case files and due diligence reports are, the more strained the GPU memory will be, and the more easily key clauses will be omitted: APS intends to adopt "topology-aware compression" to preserve the semantic structure of long texts.

氪友IBNF2026-10-09 16:39
The APS long-text compression solution has been approved by the legal department and is seeking seed round cooperation.

The Contradiction Between Video Memory Pressure of Long-Context Inference and Accuracy Requirements in Professional Scenarios

When large language models process long documents such as contracts, case files and due diligence reports, the longer the context, the higher the inference video memory and computing overhead. For scenarios such as legal affairs, finance and government administration, the real difficulty is not only "being able to run the model", but also "not losing key facts after compression". Many existing compression ideas seem effective on ordinary texts, but when applied to long contracts, complex ownership relationships and cross-paragraph fact tracing, the model is prone to missing key clauses or confusing subject relationships.

Users in such scenarios do not care about what fancy capabilities the model uses, but only three things: whether the key information is retained, whether the inference basis is correct, and whether the cost can be reduced. Therefore, long-context inference compression has evolved from an engineering optimization problem to a prerequisite for the large-scale deployment of professional AI applications.

Topology-Aware Compression Idea of APS and Verified Results

The APS project proposes a relatively underlying long-context inference compression route: it does not take "attention score" as the only basis. Instead, after the model completes context encoding, it identifies regions with obvious changes in information density from the representation layer, then retains the core context in units of blocks and reserves necessary buffer areas, so that the compressed context maintains the original semantic structural relationship as much as possible.

In terms of engineering implementation, APS adheres to three seemingly "rigorous" principles: the compressed key-value cache must be physically reorganized; each retained context fragment must carry its real position information in the original text; the position encoding and attention mask in the generation stage must be strictly aligned. The only purpose of doing so is to make the model still know "where each sentence was originally located in the document" even when half of the context is cut off, so as to avoid the long-text relationship from being destroyed by compression.

Under the experimental setting of Qwen2.5-7B, 32K context and 50% key-value cache compression, APS achieved a 9/9 result in the long-text key information retrieval task; in fifteen legal probe tests, APS performed on a par with the full uncompressed baseline, and outperformed the random cropping scheme. More critically, feedback from the scenario side shows that after the technical team of Tongyi Farui ecological service provider conducted tests with real legal corpus, they believed that APS has "absolute advantages" in key information extraction from long documents, and stated that further communication on cooperation forms can be carried out if the project party has the willingness to cooperate.

Team Background, Technology Protection and Current Progress

APS is promoted by a team with cross-background of law and artificial intelligence. The founder graduated from the Law School of Indiana University, USA, and once served as legal director of multiple Fortune Global 500 enterprises. He has long been engaged in high-frequency long-text tasks such as contract review, intellectual property rights, compliance and dispute resolution, and has direct experience of "which information cannot be lost in legal scenarios".

In terms of technology protection, the core algorithm of APS does not follow the public patent route, but is protected as a trade secret. The project party judges that in the very early stage, premature disclosure of implementation details is not conducive to the maintenance of technical barriers; whether to apply for patents in the future will be determined according to the cooperation mode, authorization scope and commercialization path.

At the present stage, APS has completed the engineering verification and preliminary evaluation of the core idea, and has not yet established a financing entity. It is seeking seed round support based on technical verification results and scenario feedback. If the funds are in place, they will be mainly used for engineering encapsulation, improvement of the evaluation system, joint verification with scenario parties and establishment of the corporate entity. What the project party values more is not simply getting a sum of money, but finding industrial parties or early-stage investors who can understand the value of long-context inference infrastructure and are willing to jointly verify with real business corpus first.

Long-context inference compression will not be "just another AI story" in the short-term hot trend. As enterprise-level applications shift from chat scenarios to real long documents such as contracts, cases, research reports and regulatory documents, inference cost and information fidelity will become the watershed for whether AI can be integrated into production systems. What APS currently demonstrates is not a completed commercial closed loop, but a technical direction recognized by scenario parties as "having absolute advantages", as well as a set of methodologies that prevent the semantic structure from collapsing as much as possible under a high compression ratio.