When modern artificial intelligence language models reason using step-by-step chain-of-thought pathways or exchange free-text intermediate outputs, they are essentially performing a translation task. They take deeply structured, hierarchical information and serialize it into standard natural language. But a fundamental, pressing question has lingered within the artificial intelligence research community: how much of that intricate, tree-structured compositional content actually survives this natural language bottleneck? To answer this question empirically, a team of researchers—Xavier Suau, Alex Ferrando de las Morenas, Luca Zappella, and Samy Bengio—has proposed and executed a novel round-trip evaluation protocol specifically designed for tree-structured expressions. Their findings offer a sobering look at how large language models handle complex information, shedding light on the hidden limits of linguistic serialization in machine reasoning. Read Also: REFACTOR-VLA Framework Introduces Wake-Sleep Architecture to Solve Long-Horizon Bottlenecks in Vision-Language-Action Models New Research Breakthrough Bridges Efficiency Gap in Federated Optimization and Stochastic Variational Inequalities The research team set up a rigorous experimental pipeline to test the fidelity of this communication channel. First, a generator model converts a procedurally generated arithmetic expression into a standard word problem. Next, a separate extractor model attempts to recover the original mathematical expression from the word problem entirely on its own, without seeing the intermediate prompt generation. To measure success precisely, the team used symbolic equivalence as an exact oracle, verifying whether the extracted expression matched the original algebraic logic down to the finest detail. By evaluating every possible pairwise combination across sixteen distinct models, the researchers constructed a comprehensive communication matrix. The marginals of this matrix successfully separated the inherent generation quality of the models from their extraction capabilities, providing a granular view of where communication breaks down. From this extensive testing, three primary findings emerged, painting a detailed picture of how models handle hierarchical data serialization. The first major finding is that the communication channel between language models using natural language is both lossy and asymmetric. When the researchers swapped which model acted as the generator and which acted as the extractor, accuracy shifted by up to 60.4 percentage points. Perhaps most counterintuitively, the highest overall accuracy—reaching 92.9%—was achieved by combining different models on each end of the pipeline rather than relying on the exact same model for both generation and extraction. This suggests that complementary strengths across different architectures can bridge some of the gaps inherent in natural language serialization, whereas using a uniform model on both ends does not automatically guarantee better transmission of structured thought. The second finding dives into the root causes of these communication failures. The researchers discovered that at least 73.6% of all round-trip failures originate strictly at the generation phase. Furthermore, the difficulty of the task is primarily driven by the underlying tree structure of the expressions—such as the total operator count, tree depth, and right-branching complexity—rather than the specific model family doing the work. As the hierarchical structures become more complex, the language models struggle proportionally more to encode those relationships into fluent prose, highlighting a structural limitation in how text represents trees. The third major finding demonstrates that this flawed communication channel is trainable. The team found that approximately 3,600 fine-tuning examples sharing the evaluation’s specific operators and tree shapes could lift every tested open-weight model above the performance of an untrained Gemini-3.1-Pro, which served as a strong upper bound under matched semantics. To ensure the performance boost was not merely a side effect of matched semantics, the researchers also tested a disjoint-domain regime introducing entirely new operators and vocabulary. This secondary test also raised the performance of every open-weight model, confirming that the improvement represents genuine adaptability, though a noticeable performance gap to the frontier models still remains. Together, these results identify tree-structured expression serialization as a primary, foundational limiting factor whenever artificial intelligence models attempt to communicate hierarchical structure through the medium of natural language. As the industry increasingly leans on chain-of-thought reasoning and multi-agent systems that pass text back and forth, understanding these bottlenecks becomes crucial for building more reliable and coherent AI architectures. Related readings and updates in the broader research landscape highlight the ongoing evolution of foundational language systems. With the ongoing integration of generative AI into consumer applications while prioritizing data privacy, major technology developers continue to refine on-device and server-based foundation models. Concurrently, lightweight syntactic analysis tools like Semgrep and Comby leverage the tree structure of code, making them significantly more expressive than traditional string and regular expression searches. Unlike conventional framework architectures that analyze codebases via explicit syntax tree manipulations, these specialized tools utilize query languages closely resembling the source language itself, though matching techniques still require careful handling of complete versus incomplete fragments. Post navigation Shared Selective Persistent Memory Architecture Boosts Agentic LLM Task Completion to 96 Percent, Study Shows