A recent study highlights that despite high agreement between large language models (LLMs) and humans on moral judgments, this does not necessarily imply that their moral grounds are aligned. The researchers created a benchmark of 500 ethics-derived items across five domains of moral judgment, with new human annotator and LLM annotations of both final labels and supporting rationales. The results show that while LLMs often achieve high agreement with human annotator majority labels, rationale-level analysis reveals systematic divergence in the moral grounds expressed by humans and models. Notably, models redistribute attention across categories such as harm, respect, promise-keeping, justice, desert, and excuse relevance, even when their final labels match the human annotator majority. This study underscores the limitation of relying solely on label-based evaluation for assessing LLMs' moral alignment and suggests that analysis of reasons, principles, and moral priorities should be incorporated into the evaluation framework. Blogger's Review: This study reveals the complexity of AI in moral judgments, emphasizing that agreement on labels is not enough. It is necessary to delve deeper into the moral foundations and decision-making processes of AI systems to truly assess their moral alignment with humans.