#Public domain text?
20-Mar-93 21:39:51
Sb: #43507-#Public domain text?
Fm: Frank/Lisa Richards 76354,15
To: Rich Bowers/OPA 71333,1114
Rich,
I never really saw the _Dataware_ full text engine in action. Their
image based stuff is really fine, but until they bought RefTech their heart
really wasn't in text. (I always assumed because they did such great page
images<g>)
I'm a little constrained talking about Reference Technology because we
are vendors to them, competitors in the data prep business, and did discuss
being customers at one time. At this late date I'm not always really sure what
I know under what nondisclosure agreement. (This of course applies to the
Cambridge end of the company too, but what's to hide about TIFF files<g>)
For our application, besides the cost, we have a real interest in
maintaining simplicity in our user interface, and according to my revered
business partner we couldn't sufficiently customize the interface.
Also, FWIW, as far as the lawyer's dream system, at last years SIGCAT,
I talked to some folks from OAG (Official Airline Guide) that had an aircraft
maintenance CDROM that actually took account of the differences from serial
number to serial number in aircraft production. This would be wonderful for
keeping track of what the law was on the date that something happened, rather
than just what it is now. Again, an idea for someone with a million bucks. (and
50 state marketing of a Federal disk to get the million back.) Even after the
NRE was paid, I'm not sure that New Hampshire or Vermont could even pay back
the on-going maintenance.
And another throwaway: We've looked into FastTag, and don't find it to
(usually) be worth the trouble. It takes some serious training time, on a
per-job basis, and you still have to have _people_ look at the document(s), and
fix Fast Tag's mistakes as well as the OCR's. For the same kind of effort, we
can have the proofreaders put in hints and then write some job- specific code
to do the tagging algorithmically.
I can see that the situation might be different if you had clean,
untagged electronic copy, from the customer or off-shore keying, or if you were
putting images on the disk and were willing to index on a dirty scan just to
get the right page image. However, the second case is dumb: If you come to us
in the first place we can deliver tagged copy for less than the cost of
offshore keying plus tagging in-house. I don't know how real the third case is:
We do get jobs that are being full text indexed with images being displayed,
but the ones that actually come to us have a higher accuracy requirement than a
dirty scan would give. It's possible that we simply don't see the ones that are
done from new 12 point laser printed copy that would OCR at 99.98% accuracy,
but when it's check-writing time we don't find people actually willing to live
with 99% accuracy.
Frank