From Advertised Improvements to Measured Capabilities: Evaluating ChatGPT Images 2.5 on Forgery Tasks
Researchers evaluated the forgery detection capabilities of the new ChatGPT Images 2.5 against the older GPT-Image-2 model. By testing the Flare and Sunburst API models on tasks with predetermined outcomes, the study aims to determine if the recently advertised performance improvements translate into measurable accuracy for identifying tampered images.
ModelsChatGPT Images 2.5
Covered by 1 source
- AarXiv CS.AI↗Ankit Raj, Yuxin Zhang, Kidus Zewde, Tommy Duong, Jiaqi Gan, Xingyu Shen, Yuchen Zhou, Huaiyu Guo, Siyu Zhang, Simiao Ren22h ago