The traditional Batak Ulos weaving industry struggles to produce diverse and innovative motifs due to manual design constraints. This paper introduces a multimodal generative framework that couples a LoRA‑fine‑tuned Stable Diffusion XL v1.0 with the multimodal LLM LLaMA 1.5‑7B, enabling controllable and culturally faithful Ulos motif generation.
The framework leverages four complementary conditioning mechanisms—text, image, latent representation, and semantic map (via ControlNet)—each governing semantic intent, visual style, latent features, and spatial layout respectively.
A five‑level ablation study across three scenarios (shape transformation, colour variation, high‑complexity input) reveals that more conditioning signals do not guarantee better results. The Text+Image+Semantic‑Map combo achieves the best FID (270) and CLIP Score (0.65‑0.70) but the lowest SSIM (0.65); Text+Image+Representation offers a balanced trade‑off with stable SSIM (0.84) and competitive FID (280); combining all four mechanisms yields the weakest FID (330), indicating conflicting optimization signals.
Qualitative evaluation with nine weavers and thirty public participants shows statistically significant higher acceptance (Wilcoxon, p=0.007), confirming the framework’s practicality.
Review