-
Open VLM Leaderboard
π1.03kVLMEvalKit Evaluation Results Collection
-
Open VLM Video Leaderboard
π135VLMEvalKit Eval Results in video understanding benchmark
-
Open LMM Reasoning Leaderboard
π₯44A Leaderboard that demonstrates LMM reasoning capabilities
-
MMBench Leaderboard
π24Explore MMBench Leaderboard data
AI & ML interests
None defined yet.
Recent Activity
View all activity
Papers
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities
CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome Reward
-
opencompass/CompassJudger-1-32B-Instruct
Text Generation β’ 33B β’ Updated β’ 75 β’ β’ 19 -
opencompass/CompassJudger-1-14B-Instruct
Text Generation β’ 15B β’ Updated β’ 14 β’ β’ 2 -
opencompass/CompassJudger-1-7B-Instruct
8B β’ Updated β’ 85 β’ 10 -
opencompass/CompassJudger-1-1.5B-Instruct
2B β’ Updated β’ 106 β’ 1
-
Open VLM Leaderboard
π1.03kVLMEvalKit Evaluation Results Collection
-
Open VLM Video Leaderboard
π135VLMEvalKit Eval Results in video understanding benchmark
-
Open LMM Reasoning Leaderboard
π₯44A Leaderboard that demonstrates LMM reasoning capabilities
-
MMBench Leaderboard
π24Explore MMBench Leaderboard data
CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome Reward
-
opencompass/CompassJudger-1-32B-Instruct
Text Generation β’ 33B β’ Updated β’ 75 β’ β’ 19 -
opencompass/CompassJudger-1-14B-Instruct
Text Generation β’ 15B β’ Updated β’ 14 β’ β’ 2 -
opencompass/CompassJudger-1-7B-Instruct
8B β’ Updated β’ 85 β’ 10 -
opencompass/CompassJudger-1-1.5B-Instruct
2B β’ Updated β’ 106 β’ 1