@hn_56ff6c
about 2 months ago
Okay results are in for GenAI Showdown with the new gpt-image 1.5 model for the editing portions of the site! https://genai-showdown.specr.net/image-editing Conclusions - OpenAI has always had some of the strongest prompt understanding alongside the weakest image fidelity. This update goes some way towards addressing this weakness. - It's leagues better at making localized edits without altering the entire image's aesthetic than gpt-image-1, doubling the previous score from 4/12 to 8/12 and the only model that legitimately passed the Giraffe prompt. - It's one of the most steerable models with a 90% compliance rate Updates to GenAI Showdown - Added outtakes sections to each model's detailed report in the Text-to-Image category, showcasing notable failures and unexpected behaviors. - New models have been added including REVE and Flux.2 Dev (a new locally hostable model). - Finally got around to implementing a weighted scoring mechanism which considers pass/fail, quality, and compliance for a more holistic model evaluation (click pass/fail icon to toggle between scoring methods). If you just want to compare gpt-image-1, gpt-image-1.5, and NB Pro at the same time: https://genai-showdown.specr.net/image-editing?models=o4,nbp
