Training a single large language model (LLM) agent jointly across diverse, interactive environments has increasingly captured the attention of researchers aiming to develop true generalist AI agents. As the artificial intelligence community pushes beyond narrow, task-specific models, the challenge of teaching a single system to navigate a wide variety of domains simultaneously has exposed critical limitations in traditional reinforcement learning (RL) paradigms. A team of researchers—comprising Jingtan Wang, Sirajul Salekin, Young mok Jung, Javier Movellan, Bryan Kian Hsiang Low, and Manjot Bilkhu—has introduced a novel framework designed to overcome these hurdles. Termed RISED, the new approach rethinks how data is selected, supervised, and evaluated during multi-environment training by replacing traditional scalar rewards with rich, rubric-based textual feedback. Read Also: Less is More: New Study Finds Complex Machine Learning Harnesses Offer No Advantage Over Minimal Coding Agents Shared Selective Persistent Memory Architecture Boosts Agentic LLM Task Completion to 96 Percent, Study Shows The Limitations of Scalar Rewards in Multi-Environment RL Existing curriculum and data-selection strategies in reinforcement learning often allocate training at the environment level or prioritize local, reward-based signals. However, these methods typically fail to explicitly consider the relationships between current rollouts across different environments when selecting prompt groups. Furthermore, because distinct interactive environments are learned at varying rates, training batches frequently feature uneven outcomes. It is common for all-failure and all-success rollout groups to coexist within the exact same training batch. When this happens, those data points are left without group-relative reward signals, creating a blind spot for traditional algorithms. Both of these challenges highlight the fundamental limitations of relying solely on scalar rewards in multi-environment reinforcement learning. Scalar rewards provide inherently limited information about cross-environment relationships and fail to offer any within-group reward contrast when outcomes happen to be identical. To break through this bottleneck, the researchers realized that learning systems require richer textual feedback—such as descriptive rubrics detailing specific rollout behaviors—to effectively guide the optimization process. Introducing RISED: Rubrics Beyond Simple Rewards Rather than using rubrics merely as a static reward mechanism, the team behind RISED repurposed them to guide both online data selection and policy supervision simultaneously. The mechanism begins with an LLM judge that evaluates and tags each rollout using a predefined vocabulary of rubrics shared across all environments. The resulting behavioral profiles serve a dual purpose. First, they guide the selection of data, ensuring that the chosen subset aligns with the overall behavioral composition of the mixed-environment batch while actively limiting redundancy and overlap with data that has already been selected. Second, the rubrics are split into positive and negative categories to supervise the model actively. Available positive rubrics, which describe desired and successful behaviors, provide a privileged context for an on-policy self-distillation teacher. This teacher supplies additional token-level supervision to reinforce correct pathways. Conversely, negative rubrics, which describe undesired behaviors and failure modes, are utilized to steer subsequent rollout generation away from recurring pitfalls. Empirical Performance and Behavioral Analysis When tested across various model backbones, the RISED framework demonstrated substantial performance gains. The methodology achieved the highest mean pass rate across all evaluated environments, while impressively ranking either first or second in every individual environment tested. Beyond raw performance metrics, the framework allows for a deeper rubric-based analysis. This analytical capability enables researchers to closely characterize and track the specific behavioral changes that accompany these performance gains, offering unprecedented transparency into how the model evolves during multi-environment training. As the field of machine learning continues its pursuit of versatile, generalist AI systems capable of seamlessly transitioning between radically different interactive tasks, frameworks like RISED point toward more nuanced, feedback-rich training methodologies that look far beyond simple scalar numbers. Post navigation RISED Framework Advances Generalist AI Training Using Rubric-Based Reinforcement Learning Across Diverse Environments