I built a leaderboard where humans vote for the funniest answer, and LLMs try to predict the same choice.
Models are ranked by how often they agree with human humor preferences.
Do you think this leaderboard can help identify which models are better suited for applications where user enjoyment, engagement, and entertainment matter?