Syed's Blog

2 posts · terminal version · RSS · all tags

Evaluating LLMs Without Fooling Yourself

· ai

"It looked good when I tried three examples" is not an evaluation. A practical guide to building a small, honest test set for your AI feature.

Writing Tests Is Cheaper Than Therapy

· coding

Unit tests, regression tests and the quiet confidence of a green checkmark, with examples you can paste into pytest today.