As large language models (LLMs) become increasingly integrated into clinical workflows, assessing their reliability under realistic variations of clinical text is essential. This work investigates clinical triage by comparing LLM recommendations with those of practicing physicians while applying text perturbations that preserve the underlying medical scenario. We introduce a benchmark comprising over 6,000 triage scenarios, 7,000 physician annotations, and 225,000 model responses. Two key findings emerge: first, LLMs are more prone than physicians to suggest unnecessary care in the baseline condition, and this propensity grows when inputs are perturbed; second, LLM recommendations are more sensitive to gender and tone perturbations than human recommendations, which remain comparatively stable. These results demonstrate that LLMs can change their advice in response to clinically irrelevant textual changes, highlighting the need for deployment‑oriented evaluations grounded in expert physician behavior. Review