We introduce CulturalMenuBench, a dataset of 4,870 items covering ten languages and eighteen regions. Each item pairs a final‑dish image and step‑by‑step cooking images with ingredient lists, procedural text, and regional labels, forming ten tasks that range from basic recognition to process‑grounded cultural attribution.\ \ Evaluating twelve multimodal language models reveals a pronounced knowledge‑application gap: while models exceed 94% accuracy on standard four‑choice questions, their performance drops to at most 56% when asked to attribute dishes to Chinese regional cuisines, despite using the same format.\ \ Diagnostic analysis shows error patterns consistent with random guessing, accuracy correlates with visual distinctiveness rather than cultural structure, and models achieve 7‑18 points higher accuracy when using dish names alone versus images, indicating that cultural knowledge exists but is not activated by visual input.\ \ An ablation study confirms that the tasks truly require procedural evidence: removing sequential cooking images selectively harms process‑grounded tasks while leaving others stable.\ \ Overall, CulturalMenuBench demonstrates that near‑perfect recognition can mask an inability to apply cultural knowledge, motivating training approaches that explicitly connect perception, procedure, and cultural context.\ \ Review