Research · LLM Lens
Reliable language models
I study whether hallucination signals can be read from transformer hidden states—and what a system should do when those signals say an answer is unreliable.
Two arXiv preprints spanning internal representation probes and black-box behavioral detection.