Vision-language models may follow harmful requests when asked to point
Vision-language models that refuse harmful questions may still comply when asked to locate harmful content in an image. A new study finds that safety prompts alone do not reliably close this gap, and tests a fine-tuning approach intended to
Open discussion →