This is cool. I've always thought LLM humor understanding is an under researched, but important, topic, especially w.r.t "is AGI here yet?".
I voted on a few "Which one is funnier?" choices, but honestly none of them were funny. The em dashes in every joke already set the "another LLM slop" mood. I would suggest re-writing the prompts to drop dashes.
Suggestion #2: introduce basic quality checks. There was this one that talked about a 2020 event as if it was still indeterminate: "The date for Superbowl 2020 has been announced as Sunday, February 2 ... They haven't yet announced who the Patriots will be playing."
Suggestion #3: slider instead of A/B single choice. Sometimes I leaned towards one choice, but it still wasn't that funny to select it; a more nuanced scale would have helped there.
Thanks for building this. From the first glance, we humans are safe from machine AGI judging by joke quality today.
hey, builder here. The short version:
LLMs take three tests:
explain why jokes work (or don't), write jokes under shared premises or predict which jokes humans prefer
The finding so far that surprised me: every model aces explaining real jokes (95%+) but drops hard on explaining why a failed joke fails (81–92%).
happy to answer anything about the eval system and open to any sort of feedback!
This is cool. I've always thought LLM humor understanding is an under researched, but important, topic, especially w.r.t "is AGI here yet?".
I voted on a few "Which one is funnier?" choices, but honestly none of them were funny. The em dashes in every joke already set the "another LLM slop" mood. I would suggest re-writing the prompts to drop dashes.
Suggestion #2: introduce basic quality checks. There was this one that talked about a 2020 event as if it was still indeterminate: "The date for Superbowl 2020 has been announced as Sunday, February 2 ... They haven't yet announced who the Patriots will be playing."
Suggestion #3: slider instead of A/B single choice. Sometimes I leaned towards one choice, but it still wasn't that funny to select it; a more nuanced scale would have helped there.
Thanks for building this. From the first glance, we humans are safe from machine AGI judging by joke quality today.
[flagged]