Abstract
Recommender systems increasingly need to choose among heterogeneous agents, including collaborative filters, sequential models, content-based retrievers, and LLM-based rerankers, yet no single agent performs best across all requests. We study this problem as task-aware agent ranking under cost constraints through RouteRec, a lightweight framework that first evaluates request-level hard selection and then tests an item-level learned aggregation relaxation.
Using MovieLens-1M with four low-cost agents and one LLM reranker, we find that oracle selection offers substantial headroom: HR@10 reaches 0.584 with all agents and 0.508 with cheap agents only, showing that useful cross-agent signal exists. However, under a leakage-free 5-fold out-of-fold protocol, hard selection fails to realise this potential. RouteRec-Select with cheap agents achieves HR@10 of 0.223 ± 0.008, below BM25 at 0.254, while LLM gating reaches only 0.215 ± 0.016.
In contrast, the same strict protocol shows more promising results for learned shortlist aggregation, which combines deployable ranks and scores at the item level. RouteRec-StackCheap matches BM25 in HR@10 and obtains a higher NDCG point estimate, 0.123 compared with 0.114. Gated all-agent aggregation further improves HR@10 to 0.295 ± 0.022 while using LLM calls for 70.2% of requests. These findings suggest that hard request-level agent selection is too coarse for this sparse offline setting, whereas item-level aggregation provides a more effective action space for recommender-agent routing.
Using MovieLens-1M with four low-cost agents and one LLM reranker, we find that oracle selection offers substantial headroom: HR@10 reaches 0.584 with all agents and 0.508 with cheap agents only, showing that useful cross-agent signal exists. However, under a leakage-free 5-fold out-of-fold protocol, hard selection fails to realise this potential. RouteRec-Select with cheap agents achieves HR@10 of 0.223 ± 0.008, below BM25 at 0.254, while LLM gating reaches only 0.215 ± 0.016.
In contrast, the same strict protocol shows more promising results for learned shortlist aggregation, which combines deployable ranks and scores at the item level. RouteRec-StackCheap matches BM25 in HR@10 and obtains a higher NDCG point estimate, 0.123 compared with 0.114. Gated all-agent aggregation further improves HR@10 to 0.295 ± 0.022 while using LLM calls for 70.2% of requests. These findings suggest that hard request-level agent selection is too coarse for this sparse offline setting, whereas item-level aggregation provides a more effective action space for recommender-agent routing.
| Original language | English |
|---|---|
| Publication status | Published - 24 Jul 2026 |
| Event | AgentSearch: The First Workshop on Indexing, Retrieval, and Ranking of AI Agents - Naarm, Australia Duration: 24 Jul 2026 → 24 Jul 2026 |
Workshop
| Workshop | AgentSearch |
|---|---|
| Country/Territory | Australia |
| City | Naarm |
| Period | 24/07/26 → 24/07/26 |
Fingerprint
Dive into the research topics of 'RouteRec: Strict Evaluation of Recommender-Agent Selection and Aggregation'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver