A. Surnames
Surnames are a big problem because
they can contain apostrophes like in O'Neil, accents like in Menaḅ, or because they can be compound like Dalla Libera or double, like Conte Camerino.
So, at BioPD we decided to index SURNAMES following the
rules reported in the following examples.
Another problem related to fields
containing SURNAMES is finding a subpopulation of people
all having the same surname and the same first letter of the
name. Suppose we want to find all people whose surname is Smith
and whose name begins with J.
At the moment at BioPD, we do not have databases with these
features, but we plan to index fields containing both the surname
and the initial of the name connecting them with a -
(minus, hyphen) because freeWAIS-sf DOES NOT allow to
index one letter word.
We also plan to index this field with stemming on. This way searches can be performed in two ways:

In this case only people called Smith J. are found.

In this case all people whose
surname stem is Smith are found.
On the contrary, if the field has full surname and name, this
rule is NOT followed, of course.
© 1996-2003 BioPD -
University of Padova (Italy) - Author: Leopoldo
Saggin - Last version: January
28, 2003
Best efforts were made to provide correct information, however
this document may contain technical inaccuracies and/or
typographical errors.
The author declares that this material is provided "as
is" without any warranty even in the implied warranty of
merchantability or fitness for a particular purpose.
All trademarks cited inside this document are property of their
respective owners