Thursday, December 27, 2007

OCR Technology - Leadorganizer.net

We are talking here Optical character recognition Technology as part of Document Management functions of leadorganizer.net

Recognition of cursive text is an active area of research, with recognition rates even lower than that of hand-printed text. Higher rates of recognition of general cursive script will likely not be possible without the use of contextual or grammatical information.

For example, recognizing entire words from a dictionary is easier than trying to parse individual characters from script. Reading the Amount line of a cheque (which is always a written-out number) is an example where using a smaller dictionary can increase recognition rates greatly. Knowledge of the grammar of the language being scanned can also help determine if a word is likely to be a verb or a noun, for example, allowing greater accuracy. The shapes of individual cursive characters themselves simply do not contain enough information to accurately (greater than 98%) recognize all handwritten cursive script.

For more complex recognition problems, intelligent character recognition systems are generally used, as artificial neural networks can be made indifferent to both affine and non-linear transformations.


ref:Insurance Document Management Software, wikipedia

Monday, December 24, 2007

OCR Technology

We are talking Document Management System here and as a part of DMS, we talked about OCR in our last post. Today we are going to talk more about OCR Technology.

Current Status of OCR Technology:

The accurate recognition of Latin-script, typewritten text is now considered largely a solved problem. Typical accuracy rates exceed 99%, although certain applications demanding even higher accuracy require human review for errors. Handwriting recognition, including recognition of hand printing, cursive handwriting, is still the subject of active research, as is recognition of printed text in other scripts (especially those with a very large number of characters)

Systems for recognizing hand-printed text on the fly have enjoyed commercial success in recent years. Among these are the input device for personal digital assistants such as those running Palm OS. The Apple Newton pioneered this technology. The algorithms used in these devices take advantage of the fact that the order, speed, and direction of individual lines segments at input are known. Also, the user can be retrained to use only specific letter shapes. These methods cannot be used in software that scans paper documents, so accurate recognition of hand-printed documents is still largely an open problem. Accuracy rates of 80% to 90% on neat, clean hand-printed characters can be achieved, but that accuracy rate still translates to dozens of errors per page, making the technology useful only in very limited applications. This variety of OCR is now commonly known in the industry as ICR, or Intelligent Character Recognition.

we continue our talk on OCR Technology in next post.

ref: Document Management Software, wikipedia

Thursday, December 20, 2007

OCR - Optical Character Recognition - Document Management

We talked Metadata as Document Management System components in our last post, today we are going to talk about OCR.

Optical character recognition, usually abbreviated to OCR, is the mechanical or electronic translation of images of handwritten, typewritten or printed text (usually captured by a scanner) into machine-editable text.

OCR is a field of research in pattern recognition, artificial intelligence and machine vision. Though academic research in the field continues, the focus on OCR has shifted to implementation of proven techniques. Optical character recognition (using optical techniques such as mirrors and lenses) and digital character recognition (using scanners and computer algorithms) were originally considered separate fields. Because very few applications survive that use true optical techniques, the OCR term has now been broadened to include digital image processing as well.


ref: document management software, wikipedia

Tuesday, December 18, 2007

Metadata - Document Management Software

We talked Data Capture, Data Indexing, Data Storage, workflow, Data integration and Metadata, Data Retrieval, Data Collaboration & Versioning as components of Document Management and lead organizer software.

We talked Metadata as Document Management System components. Today we talk more about Metadata.


Metadata is data about data. An item of metadata may describe an individual datum, or content item, or a collection of data including multiple content items.

Metadata (sometimes written 'meta data') is used to facilitate the understanding, use and management of data. The metadata required for effective data management varies with the type of data and context of use. In a library, where the data is the content of the titles stocked, metadata about a title would typically include a description of the content, the author, the publication date and the physical location.

In the context of a camera, where the data is the photographic image, metadata would typically include the date the photograph was taken and details of the camera settings. On a portable music player such as an Apple iPod, the album names, song titles and album art embedded in the music files are used to generate the artist and song listings, and are metadata. In the context of an information system, where the data is the content of the computer files, metadata about an individual data item would typically include the name of the field and its length.

Metadata about a collection of data items, a computer file, might typically include the name of the file, the type of file and the name of the data administrator.




ref: Document Organizer & Lead Organizer Software, wikipedia

Thursday, December 13, 2007

Collaboration & Versioning - Document Management Software

We are talking Document Management System and different components attach with it. We talked Data Capture, Data Indexing, Data Storage, workflow, Data integration and Metadata , Data Retrieval as components of Document Management and leadorganizer software.

Today we are going to talk about Collaboration & Versioning as part of document management systems.

Collaboration
Collaboration should be inherent in a EDMS. Documents should be capable of being retrieved by an authorized user and worked on. Access should be blocked to other users while work is being performed on the document.

Versioning
Versioning is a process by which documents are checked in or out of the document management system, allowing users to retrieve previous versions and to continue work from a selected point. Versioning is useful for documents that change over time and require updating, but it may be necessary to go back to a previous copy.



ref: Document Management Software, wikipedia

Tuesday, December 11, 2007

WorkFlow - Document Management System


We are talking Document Management System and different components attach with it. We talked Data Capture, Data Indexing, Data Storage, Data integration and Metadata , Data Retrieval as components of Document Management and leadorganizer software.


Today we are going to talk about Workflow as part of document management systems.

Workflow is a complex problem and some document management systems have a built in workflow module. There are different types of workflow. Usage depends on the environment the EDMS is applied to. Manual workflow requires a user to view the document and decide who to send it to.

Rules-based workflow allows an administrator to create a rule that dictates the flow of the document through an organization: for instance, an invoice passes through an approval process and then is routed to the accounts payable department. Dynamic rules allow for branches to be created in a workflow process. A simple example would be to enter an invoice amount and if the amount is lower than a certain set amount, it follows different routes through the organization.


ref: work flow management system, wikipedia

Thursday, December 6, 2007

Document Retrieval - Document Organizer

Document retrieval is defined as the matching of some stated user query against a set of free-text records. These records could be any type of mainly unstructured text, such as newspaper articles, real estate records or paragraphs in a manual. User queries can range from multi-sentence full descriptions of an information need to a few words.

Document retrieval is sometimes referred to as, or as a branch of, Text Retrieval. Text retrieval is a branch of information retrieval where the information is stored primarily in the form of text. The advent of full text searching made the job of the indexer redundant during the 1980s. Text databases became decentralized thanks to the personal computer and the CD-ROM. Text retrieval is a critical area of study today, since it is the fundamental basis of all internet search engines.



ref: Document Retrieval & Management Software, wikipedia