Evaluating Conditioning of a Vision-Language Model for SEM Fractography
Abstract
Classifying fracture mechanisms from scanning electron microscopy (SEM) images is an important step in failure analysis of materials. We evaluate whether written domain rules, visual exemplars, and automatically refined textual instructions can improve a vision-language model’s classification of brittle, ductile, and mixed fractures without parameter updates. Specifically, we ask whether visual classification errors can be converted into reusable textual rules. In our refinement pipeline, the classifier analyzes each image, while the optimizer receives only the predicted and reference labels and the classifier’s textual account of the visual evidence. None of the evaluated conditioning strategies produced a statistically supported aggregate improvement over the domain-informed zero-shot baseline. However, visual exemplars were substantially more likely to both correct errors and overturn correct predictions when three text-only conditions disagreed. Automated textual refinement similarly produces no validated improvement under the evaluated development protocol. These results show that conditioning can change VLM decisions without reliably improving them and suggest that indirect textual accounts of visual errors may provide insufficient feedback for refinement. We therefore identify image-aware optimization and automated directing of unstable cases for expert review as important directions for reliable training-free adaptation in scientific imaging.