Xiaorong Li 維護的數據集。PhD ,Intelligent Systems Lab Amsterdam.research on video and image retrieval.
Flickr-3.5M: A collection of 3.5 million social-tagged images.
Social20: A ground-truth set for tag-based social image retrieval.
Biconcepts2012test: A ground-truth set for retrieving bi-concepts (concept pairs) in unlabeled images.
neg4free: A set of negative examples automatically harvested from social-tagged images for 20 PASCAL VOC concepts.
4wikipedia featured articles 函數圖片(以及特征)以及對應的wiki文本??梢钥纯次恼?span style="font-family:Verdana; font-size:12px; background-color:rgb(197,197,197)">A New Approach to Cross-Modal Multimedia Retrieval,還有一批文章On the Role of Correlation and Abstraction in Cross-Modal Multimedia Retrieval不過還沒有下載鏈接 http://www.svcl.ucsd.edu/projects/crossmodal/
5
http://lms.comp.nus.edu.sg/research/NUS-WIDE.htm
To our knowledge, this is the largest real-world web image dataset comprising over 269,000 images with over 5,000 user-provided tags, and ground-truth of 81 concepts for the entire dataset. The dataset is much larger than the popularly available Corel and Caltech 101 datasets. Though some datasets comprise over 3 million images, they only have ground-truth for a small fraction of images. Our proposed NUS-WIDE dataset has the ground-truth for the entire dataset.
The dataset for the Microsoft Image Grand Challenge on Image Retrieval?
另外介紹cvpaper上的整理的數據集
Participate in Reproducible Research
Detection
PASCAL VOC 2009 dataset
Classification/Detection Competitions, Segmentation Competition, Person Layout Taster Competition datasets
LabelMe dataset
LabelMe is a web-based image annotation tool that allows researchers to label images and share the annotations with the rest of the community. If you use the database, we only ask that you contribute to it, from time to time, by using the labeling tool.
BioID Face Detection Database
1521 images with human faces, recorded under natural conditions, i.e. varying illumination and complex background. The eye positions have been set manually.
CMU/VASC & PIE Face datasetYale Face datasetCaltech
15,560 pedestrian and non-pedestrian samples (image cut-outs) and 6744 additional full images not containing pedestrians for bootstrapping. The test set contains more than 21,790 images with 56,492 pedestrian labels (fully visible or partially occluded), captured from a vehicle in urban traffic.
MIT Pedestrian dataset
CVC Pedestrian Datasets
CVC Pedestrian Datasets
CBCL Pedestrian Database
MIT Face dataset
CBCL Face Database
MIT Car dataset
CBCL Car Database
MIT Street dataset
CBCL Street Database
INRIA Person Data Set
A large set of marked up images of standing or walking people
INRIA car dataset
A set of car and non-car images taken in a parking lot nearby INRIA
INRIA horse dataset
A set of horse and non-horse images
H3D Dataset
3D skeletons and segmented regions for 1000 people in images
HRI RoadTraffic dataset
A large-scale vehicle detection dataset
BelgaLogos
10000 images of natural scenes, with 37 different logos, and 2695 logos instances, annotated with a bounding box.
FlickrBelgaLogos
10000 images of natural scenes grabbed on Flickr, with 2695 logos instances cut and pasted from the BelgaLogos dataset.
FlickrLogos-32
The dataset FlickrLogos-32 contains photos depicting logos and is meant for the evaluation of multi-class logo detection/recognition as well as logo retrieval methods on real-world images. It consists of 8240 images downloaded from Flickr.
TME Motorway Dataset
30000+ frames with vehicle rear annotation and classification (car and trucks) on motorway/highway sequences. Annotation semi-automatically generated using laser-scanner data. Distance estimation and consistent target ID over time available.
PHOS (Color Image Database for illumination invariant feature selection)
Phos is a color image database of 15 scenes captured under different illumination conditions. More particularly, every scene of the database contains 15 different images: 9 images captured under various strengths of uniform illumination, and 6 images under different degrees of non-uniform illumination. The images contain objects of different shape, color and texture and can be used for illumination invariant feature detection and selection.
CaliforniaND: An Annotated Dataset For Near-Duplicate Detection In Personal Photo Collections
California-ND contains 701 photos taken directly from a real user's personal photo collection, including many challenging non-identical near-duplicate cases, without the use of artificial image transformations. The dataset is annotated by 10 different subjects, including the photographer, regarding near duplicates.
Classification
PASCAL VOC 2009 dataset
Classification/Detection Competitions, Segmentation Competition, Person Layout Taster Competition datasets
A dataset for testing object class detection algorithms. It contains 255 test images and features five diverse shape-based classes (apple logos, bottles, giraffes, mugs, and swans).
Flower classification data sets
17 Flower Category Dataset
Animals with attributes
A dataset for Attribute Based Classification. It consists of 30475 images of 50 animals classes with six pre-extracted feature representations for each image.
Stanford Dogs Dataset
Dataset of 20,580 images of 120 dog breeds with bounding-box annotation, for fine-grained image categorization.
Recognition
Face and Gesture Recognition Working Group FGnet
Face and Gesture Recognition Working Group FGnet
Feret
Face and Gesture Recognition Working Group FGnet
PUT face
9971 images of 100 people
Labeled Faces in the Wild
A database of face photographs designed for studying the problem of unconstrained face recognition
Urban scene recognition
Traffic Lights Recognition, Lara's public benchmarks.
PubFig: Public Figures Face Database
The PubFig database is a large, real-world face dataset consisting of 58,797 images of 200 people collected from the internet. Unlike most other existing face datasets, these images are taken in completely uncontrolled situations with non-cooperative subjects.
YouTube Faces
The data set contains 3,425 videos of 1,595 different people. The shortest clip duration is 48 frames, the longest clip is 6,070 frames, and the average length of a video clip is 181.3 frames.
MSRC-12: Kinect gesture data set
The Microsoft Research Cambridge-12 Kinect gesture data set consists of sequences of human movements, represented as body-part locations, and the associated gesture to be recognized by the system.
QMUL underGround Re-IDentification (GRID) Dataset
This dataset contains 250 pedestrian image pairs + 775 additional images captured in a busy underground station for the research on person re-identification.
Person identification in TV series
Face tracks, features and shot boundaries from our latest CVPR 2013 paper. It is obtained from 6 episodes of Buffy the Vampire Slayer and 6 episodes of Big Bang Theory.
ChokePoint Dataset
ChokePoint is a video dataset designed for experiments in person identification/verification under real-world surveillance conditions. The dataset consists of 25 subjects (19 male and 6 female) in portal 1 and 29 subjects (23 male and 6 female) in portal 2.
Tracking
BIWI Walking Pedestrians dataset
Walking pedestrians in busy scenarios from a bird eye view
"Central" Pedestrian Crossing Sequences
Three pedestrian crossing sequences
Pedestrian Mobile Scene Analysis
The set was recorded in Zurich, using a pair of cameras mounted on a mobile platform. It contains 12'298 annotated pedestrians in roughly 2'000 frames.
Head tracking
BMP image sequences.
KIT AIS Dataset
Data sets for tracking?vehicles?and?people?in aerial image sequences.
MIT Traffic Data Set
MIT traffic data set is for research on activity analysis and crowded scenes. It includes a traffic video sequence of 90 minutes long. It is recorded by a stationary camera.
Segmentation
Image Segmentation with A Bounding Box Prior dataset
Ground truth database of 50 images with: Data, Segmentation, Labelling - Lasso, Labelling - Rectangle
PASCAL VOC 2009 dataset
Classification/Detection Competitions, Segmentation Competition, Person Layout Taster Competition datasets
Motion Segmentation and OBJCUT data
Cows for object segmentation, Five video sequences for motion segmentation
Geometric Context Dataset
Geometric Context Dataset: pixel labels for seven geometric classes for 300 images
Crowd Segmentation Dataset
This dataset contains videos of crowds and other high density moving objects. The videos are collected mainly from the BBC Motion Gallery and Getty Images website. The videos are shared only for the research purposes. Please consult the terms and conditions of use of these videos from the respective websites.
CMU-Cornell iCoseg Dataset
Contains hand-labelled pixel annotations for 38 groups of images, each group containing a common foreground. Approximately 17 images per group, 643 images total.
Segmentation evaluation database
200 gray level images along with ground truth segmentations
The Berkeley Segmentation Dataset and Benchmark
Image segmentation and boundary detection. Grayscale and color segmentations for 300 images, the images are divided into a training set of 200 images, and a test set of 100 images.
Weizmann horses
328 side-view color images of horses that were manually segmented. The images were randomly collected from the WWW.
Saliency-based video segmentation with sequentially updated priors
10 videos as inputs, and segmented image sequences as ground-truth
Foreground/Background
Wallflower Dataset
For evaluating background modelling algorithms
Foreground/Background Microsoft Cambridge Dataset
Foreground/Background segmentation and Stereo dataset from Microsoft Cambridge
The SABS (Stuttgart Artificial Background Subtraction) dataset is an artificial dataset for pixel-wise evaluation of background models.
Saliency Detection (source)
AIM
120 Images / 20 Observers (Neil D. B. Bruce and John K. Tsotsos 2005).
LeMeur
27 Images / 40 Observers (O. Le Meur, P. Le Callet, D. Barba and D. Thoreau 2006).
Kootstra
100 Images / 31 Observers (Kootstra, G., Nederveen, A. and de Boer, B. 2008).
DOVES
101 Images / 29 Observers (van der Linde, I., Rajashekar, U., Bovik, A.C., Cormack, L.K. 2009).
Ehinger
912 Images / 14 Observers (Krista A. Ehinger, Barbara Hidalgo-Sotelo, Antonio Torralba and Aude Oliva 2009).
NUSEF
758 Images / 75 Observers (R. Subramanian, H. Katti, N. Sebe1, M. Kankanhalli and T-S. Chua 2010).
JianLi
235 Images / 19 Observers (Jian Li, Martin D. Levine, Xiangjing An and Hangen He 2011).
Extended Complex Scene Saliency Dataset (ECSSD)
ECSSD contains 1000 natural images with complex foreground or background. For each image, the ground truth mask of salient object(s) is provided.
Video Surveillance
CAVIAR
For the CAVIAR project a number of video clips were recorded acting out the different scenarios of interest. These include people walking alone, meeting with others, window shopping, entering and exitting shops, fighting and passing out and last, but not least, leaving a package in a public place.
ViSOR
ViSOR contains a large set of multimedia data and the corresponding annotations.
Oxford reconstruction data set (building reconstruction)
Oxford colleges
Multi-View Stereo dataset (Vision Middlebury)
Temple, Dino
Multi-View Stereo for Community Photo Collections
Venus de Milo, Duomo in Pisa, Notre Dame de Paris
IS-3D Data
Dataset provided by Center for Machine Perception
CVLab dataset
CVLab dense multi-view stereo image database
3D Objects on Turntable
Objects viewed from 144 calibrated viewpoints under 3 different lighting conditions
Object Recognition in Probabilistic 3D Scenes
Images from 19 sites collected from a helicopter flying around Providence, RI. USA. The imagery contains approximately a full circle around each site.
Multiple cameras fall dataset
24 scenarios recorded with 8 IP video cameras. The first 22 first scenarios contain a fall and confounding events, the last 2 ones contain only confounding events.
Action
UCF Sports Action Dataset
This dataset consists of a set of actions collected from various sports which are typically featured on broadcast television channels such as the BBC and ESPN. The video sequences were obtained from a wide range of stock footage websites including BBC Motion gallery, and GettyImages.
UCF Aerial Action Dataset
This dataset features video sequences that were obtained using a R/C-controlled blimp equipped with an HD camera mounted on a gimbal.The collection represents a diverse pool of actions featured at different heights and aerial viewpoints. Multiple instances of each action were recorded at different flying altitudes which ranged from 400-450 feet and were performed by different actors.
UCF YouTube Action Dataset
It contains 11 action categories collected from YouTube.
UCF50 is an action recognition dataset with 50 action categories, consisting of realistic videos taken from YouTube.
ASLAN
The Action Similarity Labeling (ASLAN) Challenge.
MSR Action Recognition Datasets
The dataset was captured by a Kinect device. There are 12 dynamic American Sign Language (ASL) gestures, and 10 people. Each person performs each gesture 2-3 times.
KTH Recognition of human actions
Contains six types of human actions (walking, jogging, running, boxing, hand waving and hand clapping) performed several times by 25 subjects in four different scenarios: outdoors, outdoors with scale variation, outdoors with different clothes and indoors.
Hollywood-2 Human Actions and Scenes dataset
Hollywood-2 datset contains 12 classes of human actions and 10 classes of scenes distributed over 3669 video clips and approximately 20.1 hours of video in total.
Collective Activity Dataset
This dataset contains 5 different collective activities : crossing, walking, waiting, talking, and queueing and 44 short video sequences some of which were recorded by consumer hand-held digital camera with varying view point.
Olympic Sports Dataset
The Olympic Sports Dataset contains YouTube videos of athletes practicing different sports.
SDHA 2010
Surveillance-type videos
VIRAT Video Dataset
The dataset is designed to be realistic, natural and challenging for video surveillance domains in terms of its resolution, background clutter, diversity in scenes, and human activity/event categories than existing action recognition datasets.
HMDB: A Large Video Database for Human Motion Recognition
Collected from various sources, mostly from movies, and a small proportion from public databases, YouTube and Google videos. The dataset contains 6849 clips divided into 51 action categories, each containing a minimum of 101 clips.
Stanford 40 Actions Dataset
Dataset of 9,532 images of humans performing 40 different actions, annotated with bounding-boxes.
50Salads dataset
Fully annotated dataset of RGB-D video data and data from accelerometers attached to kitchen objects capturing 25 people preparing two mixed salads each (4.5h of annotated data). Annotated activities correspond to steps in the recipe and include phase (pre-/ core-/ post) and the ingredient acted upon.
Human pose/Expression
AFEW (Acted Facial Expressions In The Wild)/SFEW (Static Facial Expressions In The Wild)
Dynamic temporal facial expressions data corpus consisting of close to real world environment extracted from movies.
ETHZ CALVIN Dataset
Image stitching
IPM Vision Group Image Stitching datasets
Images and parameters for registeration
Medical
VIP Laparoscopic / Endoscopic Dataset
Collection of endoscopic and laparoscopic (mono/stereo) videos and images
Misc
Zurich Buildings Database
ZuBuD Image Database contains over 1005 images about Zurich city building.
Color Name Data SetsMall dataset
The mall dataset was collected from a publicly accessible webcam for crowd counting and activity profiling research.
QMUL Junction Dataset
A busy traffic dataset for research on activity analysis and behaviour understanding.
50 Salads?- fully annotated 4.5 hour dataset of RGB-D video + accelerometer data, capturing 25 people preparing two mixed salads each (Dundee University, Sebastian Stein)
ASLAN Action similarity labeling challenge?database (Orit Kliper-Gross)
Berkeley MHAD: A Comprehensive Multimodal Human Action Database?(Ferda Ofli)
BEHAVE Interacting Person Video Data with markup?(Scott Blunsden, Bob Fisher, Aroosha Laghaee)
CVBASE06: annotated sports videos?(Janez Pers)
G3D?- synchronised video, depth and skeleton data for 20 gaming actions captured with Microsoft Kinect (Victoria Bloom)
Hollywood 3D?- 650 3D action recognition in the wild videos, 14 action classes (Simon Hadfield)
Human Actions and Scenes Dataset?(Marcin Marszalek, Ivan Laptev, Cordelia Schmid)
HumanEva: Synchronized Video and Motion Capture Dataset for Evaluation of Articulated Human Motion (Brown University)
i3DPost Multi-View Human Action Datasets?(Hansung Kim)
i-LIDS video event image dataset (Imagery library for intelligent detection systems)?(Paul Hosner)
MIT CBCL Automated Mouse Behavior Recognition datasets?(Nicholas Edelman)
Retinal fundus images - Ground truth of vascular bifurcations and crossovers?(Univ of Groningen)
Spine and Cardiac data?(Digital Imaging Group of London Ontario, Shuo Li)
Univ of Central Florida - DDSM: Digital Database for Screening Mammography?(Univ of Central Florida)
VascuSynth?- 120 3D vascular tree like structures with ground truth (Mengliu Zhao, Ghassan Hamarneh)
York Cardiac MRI dataset?(Alexander Andreopoulos)
Face Databases
3D Mask Attack Database (3DMAD)?- 76500 frames of 17 persons using Kinect RGBD with eye positions (Sebastien Marcel)
Audio-visual database for face and speaker recognition?(Mobile Biometry MOBIO http://www.mobioproject.org/)
BANCA face and voice database?(Univ of Surrey)
Binghampton Univ 3D static and dynamic facial expression database?(Lijun Yin, Peter Gerhardstein and teammates)
BioID face database?(BioID group)
Biwi 3D Audiovisual Corpus of Affective Communication?- 1000 high quality, dynamic 3D scans of faces, recorded while pronouncing a set of English sentences.
CMU Facial Expression Database?(CMU/MIT)
CMU/MIT Frontal Faces?(CMU/MIT)
CMU/MIT Frontal Faces?(CMU/MIT)
CMU Pose, Illumination, and Expression (PIE) Database?(Simon Baker)
CSSE Frontal intensity and range images of faces?(Ajmal Mian)
Face Recognition Grand Challenge datasets?(FRVT - Face Recognition Vendor Test)
FaceTracer Database - 15,000 faces?(Neeraj Kumar, P. N. Belhumeur, and S. K. Nayar)
FDDB: Face Detection Data set and Benchmark - studying unconstrained face detection?(University of Massachusetts Computer Vision Laboratory)
FG-Net Aging Database of faces at different ages?(Face and Gesture Recognition Research Network)
Facial Recognition Technology (FERET) Database?(USA National Institute of Standards and Technology)
Hong Kong Face Sketch Database
Japanese Female Facial Expression (JAFFE) Database?(Michael J. Lyons)
LFW: Labeled Faces in the Wild?- unconstrained face recognition.?Re-labeled Faces in the Wild?- original images, but aligned using "deep funneling" method. (University of Massachusetts, Amherst)
Manchester Annotated Talking Face Video Dataset?(Timothy Cootes)
MIT Collation of Face Databases?(Ethan Meyers)
MORPH (Craniofacial Longitudinal Morphological Face Database)?(University of North Carolina Wilmington)
MIT CBCL Face Recognition Database?(Center for Biological and Computational Learning)
NIST mugshot identification database?(USA National Institute of Standards and Technology)
ORL face database: 40 people with 10 views?(ATT Cambridge Labs)
Columbia COIL-100 3D object multiple views?(Columbia University)
Densely sampled object views: 2500 views of 2 objects, eg for view-based recognition and modeling?(Gabriele Peters, Universiteit Dortmund)
German Traffic Sign Detection Benchmark?(Ruhr-Universitat Bochum)
GRAZ-02 Database (Bikes, cars, people)?(A. Pinz)
Linkoping 3D Object Pose Estimation Database?(Fredrik Viksten and Per-Erik Forssen)
Microsoft Object Class Recognition image databases?(Antonio Criminisi, Pushmeet Kohli, Tom Minka, Carsten Rother, Toby Sharp, Jamie Shotton, John Winn)
Microsoft salient object databases (labeled by bounding boxes)?(Liu, Sun Zheng, Tang, Shum)
MIT CBCL Car Data?(Center for Biological and Computational Learning)
MIT CBCL StreetScenes Challenge Framework:?(Stan Bileschi)
NEC Toy animal object recognition or categorization database?(Hossein Mobahi)
Berkeley Segmentation Dataset and Benchmark?(David Martin and Charless Fowlkes)
GrabCut Image database?(C. Rother, V. Kolmogorov, A. Blake, M. Brown)
LabelMe images database and online annotation tool?(Bryan Russell, Antonio Torralba, Kevin Murphy, William Freeman)
Surveillance
AVSS07: Advanced Video and Signal based Surveillance 2007 datasets?(Andrea Cavallaro)
ETISEO Video Surveillance Download Datasets?(INRIA Orion Team and others)
Heriot Watt Summary of datasets for human tracking and surveillance?(Zsolt Husz)
SPEVI: Surveillance Performance EValuation Initiative?(Queen Mary University London)
Udine Trajectory-based anomalous event detection dataset?- synthetic trajectory datasets with outliers (Univ of Udine Artificial Vision and Real Time Systems Laboratory)
Textures
Color texture images by category?(textures.forrest.cz)
Columbia-Utrecht Reflectance and Texture Database?(Columbia & Utrecht Universities)
DynTex: Dynamic texture database?(Renaud Piteri, Mark Huiskes and Sandor Fazekas)
Oulu Texture Database?(Oulu University)
Prague Texture Segmentation Data Generator and Benchmark?(Mikes, Haindl)
Uppsala texture dataset of surfaces and materials?- fabrics, grains, etc.
Vision Texture?(MIT Media Lab)
General Videos
Large scale YouTube video dataset?- 156,823 videos (2,907,447 keyframes) crawled from YouTube videos (Yi Yang)
Other Collections
CANTATA Video and Image Database Index site?(Multitel)
Computer Vision Homepage list of test image databases?(Carnegie Mellon Univ)
ETHZ various, including 3D head pose, shape classes, pedestrians, pedestrians, buildings?(ETH Zurich, Computer Vision Lab)
Leibe's Collection of people/vehicle/object databases?(Bastian Leibe)
Lotus Hill Image Database Collection with Ground Truth?(Sealeen Ren, Benjamin Yao, Michael Yang)
Oxford Misc, including Buffy, Flowers, TV characters, Buildings, etc?(Oxford Visual geometry Group)
PEIPA Image Database Summary?(Pilot European Image Processing Archive)
Univ of Bern databases on handwriting, online documents, string edit and graph matching?(Univ of Bern, Computer Vision and Artificial Intelligence)
The Open Video Project?(Gary Marchionini, Barbara M. Wildemuth, Gary Geisler, Yaxiao Song)
Pics 'n' Trails - Dataset of Continuously archived GPS and digital photos?(Gamhewage Chaminda de Silva)
PRINTART: Artistic images of prints of well known paintings, including detail annotations. A benchmark for automatic annotation and retrieval tasks with this database was published at ECCV. (Nuno Miguel Pinho da Silva)
RAWSEEDS SLAM benchmark datasets?(Rawseeds Project)
Robotic 3D Scan Repository?- 3D point clouds from robotic experiments of scenes (Osnabruck and Jacobs Universities)
ROMA (ROad MArkings) : Image database for the evaluation of road markings extraction algorithms?(Jean-Philippe Tarel, et al)
Stuttgart Range Image Database?- 66 views of 45 objects
UCL Ground Truth Optical Flow Dataset?(Oisin Mac Aodha)
Univ of Genoa Datasets for disparity and optic flow evaluation?(Manuela Chessa)
Validation and Verification of Neural Network Systems?(Francesco Vivarelli)
VSD: Technicolor Violent Scenes Dataset?- a collection of ground-truth files based on the extraction of violent events in movies
WILD: Weather and Illumunation Database?(S. Narasimhan, C. Wang. S. Nayar, D. Stolyarov, K. Garg, Y. Schechner, H. Peri)