Scanning & OCR Services
Newgen is a strong and proven brand in the conversion services industry. We have the required size, track record, and financial soundness to draw on our core competencies and deliver a world-class service in scanning and OCR services.
We are flexible in our approach to conversion workflows and agile in our adoption and creation of technology geared specifically to our Scanning and the Image to PDF conversion requirements, especially when we deal with other languages.
Scanning
If the input is hardcopy, the hardcopies will be checked against the incoming notification manifest for completeness. Bibliographic and physical details (title, format, extent, etc.) are entered into the project-tracking database. Materials are then bar-coded (to have a unique tracking system, and this will be used to no loss of any material), packaged, and stored in the data unit’s library until required. Flags are set for materials that require “white glove” treatment. Scanning mode (destructive or non-destructive) will be picked up from the incoming notification. The quality assurance team will perform quality checks on the scanned output to ensure that it meets the client's requirements.
OCR
Optical character recognition (OCR) is the electronic conversion and identification of scanned images, such as those from paper documents, into searchable, digital text that is readable in a PDF format. Newgen’s experts can help determine if OCR is right for legacy projects.
Newgen is fully equipped to handle all types of scanning projects including large-scale as well as the many types of paper-based projects that require special or white-glove handling to scan, such as old and rare materials. We are experienced with both destructive and non-destructive materials.
Data Entry
In addition to handling large-scale and paper-based projects that require special attention, Newgen has all the necessary tools to handle all types of text extraction processes from scanned outputs (TIFF, JPG, etc).
Assessment: Printed material is assessed for suitability for scanning/OCR. OCR effectiveness is reduced for small type, digitally printed images, non–Latin alphabet content, and ornate scripts. Workflow will also depend on the accuracy required for the output (99.5% to 99.995%). After assessment, the material is assigned to the appropriate workflow.
OCR: Each OCR engine will have known strengths (measured in terms of accuracy of rendering particular types of content—e.g., numbers, accented characters, small type); when the outputs of different engines are compared and discrepancies are found, the discrepancies will be resolved according to the weight assigned to each engine. The operator/proofreader will review and resolve discrepancies that cannot be solved in this way. A spell-check will be run to further improve quality.
OCR plus keyboarding: OCR process as above. A second data file is prepared by a keyboarder. The outputs of the OCR and keyboarding processes are compared, and discrepancies are resolved by a proofreader with reference to the printed input material.
Double keyboarding: Two keyboarders prepare data files for the complete book. These are compared by two proofreaders working independently, who will resolve discrepancies between the files with reference to the printed input material. A third proofreader will compare the outputs of the first two proofreaders to ensure that the same decisions have been taken in each case of discrepancy between the keyboarded files. The third, senior proofreader, will arbitrate in the case of different decisions having been taken. In the jargon, this is known as double-key and double-compare. (Double-key, single-compare uses only one proofreader to directly resolve discrepancies between the two keyboarded files.)
The output from OCR/keyboarding/compare is spot-checked visually against the input material by the QC. If the level of errors in the random sample is above the threshold for the level of accuracy agreed with the client, the whole batch is rejected, and the process is started over.
JATS, BITS, DITA, S1000D, Docbook, PRISM, NITF
Newgen has been in the publishing business for years, and we bring this experience to our work with XML (Extensible Markup Language) too.
We design XML-based customer-specific DTDs, review DTD rules, recommend DTD editing, interpret schemas, map DTDs and schemas, and perform XML structuring, styling, parsing, validating, and online content checks.
XML workflows are the backbone of all our XML projects. This workflow provides a user-friendly environment for seamless transitions between content structures, resulting in a feasible customer-centric solution.
We use XSLT stylesheets and in-house tools to generate visually appealing, cross-browser-compatible HTML content.
XML structuring is an integral part of our training program that helps our workforce better understand XML structure and presentation.
Newgen has implemented an exclusive XML workflow for clients to deliver print, online and ePub files simultaneously. Here are the few XML services that we provide to our clients:
DTD Authoring
Newgen supports publishers in writing DTDs (XML Document Type Declaration) to publish their content online using their preferred styles.
Newgen ensures that the DTD writing is effective by maintaining consistency in form, function and style of writing, and using structured authoring to define and enforce content structure.
The DTD defines all the general rules related to proper nesting of tags, balancing of opening and closing tags and empty tags for an XML document to be certified as valid. It specifies which tags it uses, what attributes those tags can contain, and which tags can occur inside other tags. The DTDs also describe the elements that are optional or mandatory, the order in which they can appear, their attributes and default values.
Online QA & Approval
Newgen has a dedicated Data Controller team of QA professionals who are specially trained to perform quality control checks for XML conversion projects. The Data Controllers have hands-on experience with DTD(s) and Schemas. The Data Controller team additionally checks each XML for the following:
Uploading and Online QA
25-Aug-21
23-Sep-21
27-Jun-21
09-Jun-21
09-Jun-21
23-Feb-21
20-May-20
20-May-20
15-Jun-18
08-Dec-20
21-Jun-19
20-Jun-2019
21-05-18
05-12-2018
05-02-2018
04-10-2018
13-12-2017
12-12-2017
28-11-2017
11-03-2017
28-07-17
06-02-2017
05-04-2017
04-10-2017
04-06-2017
05-11-2018
05-09-2018
05-08-2018
16-04-18
04-05-2018
22-02-18
15-01-18
01-02-2018
15-12-17
12-05-2017
12-01-2017
27-09-17
15-09-17
07-04-2017
28-06-17
24-05-17
16-05-17
14-04-17
04-12-2017