ModelTestingBed / README.md
SerialKicked's picture
Update README.md
216a944 verified
|
Raw
History Blame Contribute Delete
1.79 kB
metadata
language:
  - en
tags:
  - testing
  - llm
  - rp
  - discussion

Why? What? TL;DR?

Various tests on, well, various LLM.

Available Tests

LLM Drawing Test (2026 ongoing)

This test is meant to evaluate models in a difficult task requiring a competency in spatial awareness, image recognition, tools calls, and creativity. In a new empty chat, the model is asked to copy the provided image to the best of its abilities. Model is then evaluated on accuracy of function calls (fail rate) and the resemblance to the original drawing.

DoggoEval (Done)

The goal of this test, featuring a dog (Rex) and his owner (EsKa), is to determine if a model is good at obeying a system prompt and character card. The trick being that dogs can't talk, but LLM love to.

Limitations

I'm testing for things I'm interested in. I do not pretend any of this is very scientific or accurate: as much as I try to reduce the amount of variables, a small LLM is still a small LLM at the end of the day. The results for other seeds, or with the smallest of change, are bound to give very different results.

I usually give the different models I'm testing a fair shake in a more casual settings. I regen tons of outputs with random seeds, and while there are (large) variations, it tends to even out to the results shown in testing. Otherwise I'll make a note of it.