Financial NLP typically follows a two‑step workflow: first validate a sentiment tool against human labels to establish construct validity, then assume the tool can extract market signals. This assumes the two evaluations measure the same construct. We test this assumption in a setting where both can be observed simultaneously: a corpus of securities class‑action filings from 2002‑2025, linking 70,500 X‑messages to abnormal stock returns, with a single‑annotator human‑labelled gold sample. Five sentiment instruments (VADER, Loughran‑McDonald, FinBERT, Twitter‑RoBERTa, and an LLM annotator) are run through an identical pipeline. The findings show that the relationship between construct and predictive validity depends on the sampling convention and score representation. Under conventional instrument‑specific sampling, human agreement aligns more closely with same‑day associations than with one‑day‑ahead predictions. On a fixed‑N panel, graded rank correlations are similar across horizons, yet coarse ordering remains weak. Thus benchmark agreement confirms semantic validity but does not by itself determine predictive rankings. Moreover, 17.6% of the messages are spam, and message volume predicts neither market damage nor settlement size.
Review