NeFut Logo NeFut
中 Admin Login

[CS.AI] From Text Decisions to Pixels: A Study of Jev-Style Visual Choice Model

Published at: 2026-09-25 22:00 Last updated: 2026-09-28 00:49
#AI #Machine Learning

PixelJev is a native‑image decision interface that maps an image, a task instruction, and a runtime candidate set to a structured choice and candidate‑conditioned probabilities using compact open multimodal models. Its first implementation unifies image recognition and multiple‑choice visual question answering through an existing language‑model readout, while offering separately evaluated options for frozen inference, language‑side adaptation, and held‑out calibration.\ \ Across seven benchmark evaluations, 64‑shot source adaptation lifts Pets accuracy from 60.13% to 92.40% across optimization seeds and transfers to natural resampling, new texture labels, and A‑OKVQA without target fitting. Frozen inference already supports both VQA tasks. A prompt‑only follow‑up on Pets and ScienceQA attributes the large Pets gain to adaptation and finds a narrower output‑validity benefit of candidate readout in adapted VQA. Specialist DINOv2 probes remain stronger on source recognition, frozen 4B outperforms adapted 2B on DTD and ScienceQA, and accuracy gains do not guarantee calibrated target probabilities.\ \ These findings establish a practical starting point for general‑purpose visual decision models and highlight remaining requirements: schema robustness, cross‑family transfer, and reliable use of visual evidence.\ \ Review

Original Source: https://arxiv.org/abs/2609.29283

[h] Back to Home