jameswniu/self-hosted-llm-evals-lab

Evaluation framework for self-hosted LLMs. Systematic prompt ablation (baseline, CoT, few-shot, self-consistency voting) on Llama 3.1 8B via lm-evaluation-harness, with Wilson CI statistical analysis, determinism validation, and load testing under concurrency. Found chain-of-thought degrades accuracy 25pp at small scale.

23
/ 100
Experimental
No Package No Dependents
Maintenance 13 / 25
Adoption 1 / 25
Maturity 9 / 25
Community 0 / 25

How are scores calculated?

Stars

1

Forks

Language

Python

License

MIT

Last pushed

Mar 09, 2026

Commits (30d)

0

Get this data via API

curl "https://pt-edge.onrender.com/api/v1/quality/prompt-engineering/jameswniu/self-hosted-llm-evals-lab"

Open to everyone — 100 requests/day, no key needed. Get a free key for 1,000/day.