NEW YORK: Humans can reliably identify a cat, butterfly, car or other objects and living creatures even when they appear as nothing more than a rough silhouette. For artificial intelligence (AI), however, this remains an impossible task.
Modern image-recognition programmes perform significantly worse than human observers when it comes to perceiving overall shapes and outlines, according to a research team led at the New York University Grossman School of Medicine, writing in the journal iScience.
The root cause of the weakness lies in the fundamentally different way algorithms function. While the human brain prioritises overall shape, even modern deep neural networks cling primarily to local details, surface patterns and textures. Minor image distortions or the absence of typical details are enough to throw the systems off.
Published in September, the findings undermine plans to implement AI in self-driving cars and robotics, as well as other areas.
More than 200 AI models tested
For the study, the team led by lead author Mugihiko Kato and principal investigator Biyu J He compared the visual performance of human test subjects with that of more than 200 deep neural networks of widely varying architectures and training methods.
The researchers systematically manipulated 240 images from 48 everyday categories. These included pure black silhouettes with no internal structure and isolated patterns such as fur textures without any discernible outline.
The team also used images in which many small objects were arranged edge to edge across the entire frame, preserving internal details while destroying the typical overall contour.
None of the 200 AI models was able to fully replicate the human recognition pattern across all conditions. Whenever identification depended solely on perceiving the overall silhouette, the computer algorithms consistently fell short of the human test subjects.
Crosses instead of cats
The discrepancy was particularly stark in a follow-up experiment in which overall shape was decoupled from local detail. The researchers filled the outlines of objects and animals with many small crosses.
While humans could still readily identify the subject from its outer contour, most AI models failed entirely. Silhouettes of cats or butterflies were suddenly and repeatedly classified as "crossword puzzles", "window screen" or "chain mail", because the small cross patterns completely dominated the algorithms' analysis.
The authors stressed that today's image-recognition models are by no means as human-like as is often assumed – even though they have been trained on vast quantities of photographs taken by humans.
Modern systems trained simultaneously on images and accompanying text do achieve considerably more human-like accuracy rates overall than conventional image-recognition networks. Yet even these systems lost their advantage as soon as the overall shape was disrupted or assessed in isolation.
Significance for autonomous driving
The findings are of central importance for practical applications of AI: If these systems are meant to operate reliably in robotics, autonomous driving, or visual prosthetics and brain-computer interfaces for people with visual impairments, they must be capable of making dependable judgements in conditions of fog, low light or partial occlusion.
In such situations, humans instinctively fall back on the rough silhouette. The study provides valuable pointers as to how algorithms must be trained in future in order to develop a more holistic and therefore more robust form of vision, the team writes. – dpa
