|
D. Perez, L. Tarazon, N. Serrano, F.M. Castro, Oriol Ramos Terrades and A. Juan. 2009. The GERMANA Database. 10th International Conference on Document Analysis and Recognition.301–305.
Abstract: A new handwritten text database, GERMANA, is presented to facilitate empirical comparison of different approaches to text line extraction and off-line handwriting recognition. GERMANA is the result of digitising and annotating a 764-page Spanish manuscript from 1891, in which most pages only contain nearly calligraphed text written on ruled sheets of well-separated lines. To our knowledge, it is the first publicly available database for handwriting research, mostly written in Spanish and comparable in size to standard databases. Due to its sequential book structure, it is also well-suited for realistic assessment of interactive handwriting recognition systems. To provide baseline results for reference in future studies, empirical results are also reported, using standard techniques and tools for preprocessing, feature extraction, HMM-based image modelling, and language modelling.
|
|
|
L.Tarazon and 6 others. 2009. Confidence Measures for Error Correction in Interactive Transcription of Handwritten Text. 15th International Conference on Image Analysis and Processing. Springer Berlin Heidelberg, 567–574. (LNCS.)
Abstract: An effective approach to transcribe old text documents is to follow an interactive-predictive paradigm in which both, the system is guided by the human supervisor, and the supervisor is assisted by the system to complete the transcription task as efficiently as possible. In this paper, we focus on a particular system prototype called GIDOC, which can be seen as a first attempt to provide user-friendly, integrated support for interactive-predictive page layout analysis, text line detection and handwritten text transcription. More specifically, we focus on the handwriting recognition part of GIDOC, for which we propose the use of confidence measures to guide the human supervisor in locating possible system errors and deciding how to proceed. Empirical results are reported on two datasets showing that a word error rate not larger than a 10% can be achieved by only checking the 32% of words that are recognised with less confidence.
|
|
|
H. Chouaib, Oriol Ramos Terrades, Salvatore Tabbone, F. Cloppet and N. Vincent. 2008. Feature Selection Combining Genetic Algorithm and Adaboost Classifiers. 19th International Conference on Pattern Recognition.1–4.
|
|
|
T.O. Nguyen, Salvatore Tabbone and Oriol Ramos Terrades. 2008. Symbol Descriptor Based on Shape Context and Vector Model of Information Retrieval. Proceedings of the 8th IAPR International Workshop on Document Analysis Systems,.191–197.
|
|
|
H. Chouaib, Salvatore Tabbone, Oriol Ramos Terrades, F. Cloppet, N. Vincent and A.T. Thierry Paquet. 2008. Sélection de Caractéristiques à partir d'un algorithme génétique et d'une combinaison de classifieurs Adaboost. Colloque International Francophone sur l'Ecrit et le Document.181–186.
|
|
|
T.O. Nguyen, Salvatore Tabbone, Oriol Ramos Terrades and A.T. Thierry. 2008. Proposition d'un descripteur de formes et du modèle vectoriel pour la recherche de symboles. Colloque International Francophone sur l'Ecrit et le Document.79–84.
|
|
|
Salvatore Tabbone, Oriol Ramos Terrades and S. Barrat. 2008. Histogram of radon transform. A useful descriptor for shape retrieval. 19th International Conference on Pattern Recognition.1–4.
|
|
|
M. Visani, V.C.Kieu, Alicia Fornes and N.Journet. 2013. The ICDAR 2013 Music Scores Competition: Staff Removal. 12th International Conference on Document Analysis and Recognition.1439–1443.
Abstract: The first competition on music scores that was organized at ICDAR in 2011 awoke the interest of researchers, who participated both at staff removal and writer identification tasks. In this second edition, we focus on the staff removal task and simulate a real case scenario: old music scores. For this purpose, we have generated a new set of images using two kinds of degradations: local noise and 3D distortions. This paper describes the dataset, distortion methods, evaluation metrics, the participant's methods and the obtained results.
|
|
|
Marçal Rusiñol, V. Poulain d'Andecy, Dimosthenis Karatzas and Josep Llados. 2013. Classification of Administrative Document Images by Logo Identification. 10th IAPR International Workshop on Graphics Recognition.
Abstract: This paper is focused on the categorization of administrative document images (such as invoices) based on the recognition of the supplier's graphical logo. Two different methods are proposed, the first one uses a bag-of-visual-words model whereas the second one tries to locate logo images described by the blurred shape model descriptor within documents by a sliding-window technique. Preliminar results are reported with a dataset of real administrative documents.
|
|
|
Marçal Rusiñol, Dimosthenis Karatzas and Josep Llados. 2013. Spotting Graphical Symbols in Camera-Acquired Documents in Real Time. 10th IAPR International Workshop on Graphics Recognition.
Abstract: In this paper we present a system devoted to spot graphical symbols in camera-acquired document images. The system is based on the extraction and further matching of ORB compact local features computed over interest key-points. Then, the FLANN indexing framework based on approximate nearest neighbor search allows to efficiently match local descriptors between the captured scene and the graphical models. Finally, the RANSAC algorithm is used in order to compute the homography between the spotted symbol and its appearance in the document image. The proposed approach is efficient and is able to work in real time.
|
|