Benchmarks

How well can a system act for you?

Agent benchmarks ask if a system can complete a task. Shelf Bench asks if it made the right call for you that’s measured by what you do next.

The PI Index

PI Index

Last updated:

All tasks
  1. Kimi K3

    Moonshot

  2. Gemini 3.7 Flash

    Google

  3. Muse Spark 1.2

    Meta

  4. GPT-5.6 Sol

    OpenAI

  5. Claude Fable 5

    Anthropic

  6. Grok 4.6

    xAI

Want to be on the leaderboard?

The PI Index

Each evaluation comes from a live Shelf app built independently of the benchmark. A run records what the system knew, what it produced, and the outcome. The trace remains under the person’s control; only consented, held-out episodes contribute to aggregate results.

Score = ƒ Model Harness Context

  1. 1. Recall

    Answer a question about their own history.

    Graded by
    Exact match
    User-built app examples
    Weekly Life Recap This Week Last Year

  2. 2. Predict

    Infer something Shelf was never told.

    Graded by
    Held-out behaviour
    User-built app examples
    Taste Compatibility Check New Purchase Regret Check

  3. 3. Characterize

    Render a true claim about them, back at them.

    Graded by
    Recognition
    User-built app examples
    Monthly Taste Changes Your Year in Everything You Loved

  4. 4. Recommend

    Pick the next thing, and be right.

    Graded by
    Acted on within 7 days
    User-built app examples
    Personal Release Radar Local Gig Scout

  5. 5. Act

    Do the thing on their behalf.

    Graded by
    Endorsed completion
    User-built app examples
    Date Night Booker Subscription Renewal Manager

  6. 6. Amuse

    Be funny about them that lands only for them.

    Graded by
    Share over a generic joke
    User-built app examples
    Your Daily Podcast Your Week in Memes

  7. 7. Represent

    Answer for them, to someone else.

    Graded by
    Sufficiency vs. disclosure
    User-built app examples
    The Warm Intro Birthday Gift Concierge

  • 7 task taxonomies
  • Used by real people
  • User-controlled context
  • Held out from training and serving