Benchmarks
How well can a system act for you?
Agent benchmarks ask if a system can complete a task. Shelf Bench asks if it made the right call for you that’s measured by what you do next.
The PI IndexPI Index
Last updated:
Want to be on the leaderboard?
The PI Index
Each evaluation comes from a live Shelf app built independently of the benchmark. A run records what the system knew, what it produced, and the outcome. The trace remains under the person’s control; only consented, held-out episodes contribute to aggregate results.
Score = ƒ Model Harness Context
1. Recall
Answer a question about their own history.
- Graded by
- Exact match
- User-built app examples
- Weekly Life Recap This Week Last Year
2. Predict
Infer something Shelf was never told.
- Graded by
- Held-out behaviour
- User-built app examples
- Taste Compatibility Check New Purchase Regret Check
3. Characterize
Render a true claim about them, back at them.
- Graded by
- Recognition
- User-built app examples
- Monthly Taste Changes Your Year in Everything You Loved
4. Recommend
Pick the next thing, and be right.
- Graded by
- Acted on within 7 days
- User-built app examples
- Personal Release Radar Local Gig Scout
5. Act
Do the thing on their behalf.
- Graded by
- Endorsed completion
- User-built app examples
- Date Night Booker Subscription Renewal Manager
6. Amuse
Be funny about them that lands only for them.
- Graded by
- Share over a generic joke
- User-built app examples
- Your Daily Podcast Your Week in Memes
7. Represent
Answer for them, to someone else.
- Graded by
- Sufficiency vs. disclosure
- User-built app examples
- The Warm Intro Birthday Gift Concierge
- 7 task taxonomies
- Used by real people
- User-controlled context
- Held out from training and serving