LongMemEval-S
| System | Score | Standing |
|---|---|---|
| ContextStreamYou are here | 90.0% | Statistical parity |
| Zep | 90.2% | Inside our 95% CI |
| supermemory | 85.4% | 4.6 pt behind |
| ByteRover | 92.8%† | Different judge |
StandingMatches Zep within the 95% confidence interval and leads supermemory by 4.6 points.
† ByteRover used a Gemini judge, so its reported score is shown but not ranked against the official GPT-4o results. Competitor scores are vendor-published.