Datasets
The datasets that defined scene text detection.
What each one tests, how it is scored, and where to download it. The datasets guide compares them in one table and lists the mistakes that make results incomparable.
ICDAR 2013
The "focused" scene text benchmark. Well-framed, mostly horizontal words photographed on purpose. Where CTPN reported 0.88 F.
229 training, 233 test
2015ICDAR 2015
The "incidental" scene text benchmark. Small, blurred, rotated words captured by a wearable camera without aiming. The number most detector papers lead with.
1,000 training, 500 test
2017Total-Text
The first widely used curved-text benchmark. Polygons around words on signage, logos and packaging that bend, arc and wave.
1,255 training, 300 test
2017SCUT-CTW1500
Curved text at line level, annotated with 14-point polygons. The companion benchmark to Total-Text, with Chinese as well as English.
1,000 training, 500 test
2016COCO-Text
Text annotations layered on MS COCO photographs. Large, incidental, and full of attributes (legible or not, printed or handwritten, language).
63,686 images from MS COCO 2014
2016SynthText
Around 800,000 synthetic images with about 8 million rendered words, generated by pasting text into natural scenes with geometry-aware placement. The pretraining set almost every detector starts from.
About 800,000 synthetic images
2019ICDAR MLT 2019
Multilingual scene text across ten languages and seven scripts, the benchmark for detectors that must not assume Latin letters.
10,000 training, 10,000 test (MLT 2019); MLT 2017 had 7,200 training, 1,800 validation and 9,000 test images