Monday, 16 April 2012

Preparing to launch the Mercurius Croaticus




Mercurius Croaticus, a prosopographical and bibliographical collection of data on Croatian Latin writers, their works, editions, and manuscripts, is nearing its launching point.

The starting dataset comprises information on:

  • 269 authors

  • 1808 works

  • 5867 printed editions

  • 58 manuscripts (more to be added soon)

  • digitised copies, where available

Wednesday, 1 February 2012

Beneša's trigrams

Here is an experimental page for researching trigrams in the De morte Christi, a neo-Latin epic by Damianus Benessa (Damjan Beneša).

What did we do:
1. using a concordance program (our reliable AntConc), we found trigrams in Beneša's Latin text, which we obtained courtesy of our colleague Vlado Rezar, Beneša's modern editor

2. we reformatted the trigrams slightly, using tr and sed, to make use of the excellent PhiloLogic crapser function (it is hard not to laugh thinking about this function, because in Croatian "serem" means "I crap")

3. using curl and a simple bash script, we sent the trigrams to CroALa

4. using again sed, we filtered out the successful hits, i. e. those which produced results

5. with some more sed, the hits were turned into searches on the page linked to at the beginning: [X]. There you'll find the trigram which produced the hit, the link to a saved search, and a report on the number of occurrences found in CroALa.

Most interesting findings for us are occurrences from Marko Marulić and Jakov Bunić, close contemporaries of Beneša; Marulić and Bunić also wrote Biblical epic poems in Latin (and Marulić's epic remained in manuscript until the 1950's).

The useful sed snippet which produces the regex line, and a line immediately before it, is here:
sed -n '/Your search found/{x;1!p;g;$N;p;};h' ben-filename


(Adapted from that goldmine, the Sed one-liners.)

Thinking about PND

An important part of our research is finding the Personennamendatei (PND) number of Croatian Latin authors and adding the number to our personal data record of the author. So far, 83 authors (of 241 from our experimental set) have been connected with their PND-Nrs.

Now we're looking into the ways Wikipedia (at least, German Wikipedia) explores the PND to uniquely identify persons and connect data on them. The BEACON format seems a nice start for a small catalogue like ours. And, of course, it would be nice if Croatian Wikipedia decided to adopt something similar to the PND scheme.

Sunday, 8 January 2012

Relations are interesting

Relations are interesting. A single fact is not interesting. If it remains single, that is. We have to relate something to it to get it going.
A single table is moderately interesting. It invites us at first to build resumes and reports (discarding individual facts in the process; this is, by the way, deeply unsatisfying), and then to compare facts in proximity to each other.
Relations of multiple tables are very interesting.
Beyond that -- wait, what lies beyond that?
So many relations, branches and joins between tables that human beings cannot hold it all in their heads, I guess.

Saturday, 7 January 2012

In the neighbourhood of Dubrovnik

One of many nice capabilities of PhiloLogic text search system is the collocation table. It shows which words occur most often in the proximity of our search string.

So we gave PhiloLogic an interesting problem. There are many Latin names for the city of Dubrovnik, and even more ways to write these. We wanted to find all of them with a single search, and to see the words which co-occur with all these names.

The search string was:
epI[dt]aUr.*|rh?a[gc]Us.*|dUbr.*. 

(capital U's and I's to find u and v, i ans y)

The result is [here], nicely shortened by Google Shortener into goo.gl/PofAI.

What do we learn from the search? That Dubrovnik is an urbs and a ciuitas, that it has principatus and nobiles and senatus. Not much surprise here.

The interesting move is to compare Dubrovnik with Split (Spalatum: [X]). It can be seen at once that there ecclesia and archiepiscopus feature much more prominently.

And so on, until the map is complete.

Thursday, 5 January 2012

Rare and Medium

This wintry afternoon I followed in the footsteps of William Whitaker, the author of WORDS Latin dictionary. The program contains a list of Latin words with very precise lexicographic descriptions -- data on period, area of application, frequency etc. The last part interested me most.

Whitaker, about whom I know almost nothing, but I'd like to know more (he seems to be outside the academia) [1], was very modest and careful in his claims, repeatedly warning users of the program that its philological expertise is limited, that he relied on other authorities and sources, that the program is intended just to be a reading help, not a research tool. Nevertheless, he has produced, I believe, the most informative freely available digital reference work on Latin usage. I'd like to see a review of his work in some scholarly journal, I think he has deserved it.

Anyway, in the documentation on word frequencies Whitaker says:

FREQ guessed from the relative number of citations given by sources need not be valid, but seems to work. (...)

type FREQUENCY_TYPE is ( -- For dictionary entries
X, -- -- Unknown or unspecified
A, -- very freq -- Very frequent, in all Elementary Latin books, top 1000+ words
B, -- frequent -- Frequent, next 2000+ words
C, -- common -- For Dictionary, in top 10,000 words
D, -- lesser -- For Dictionary, in top 20,000 words
E, -- uncommon -- 2 or 3 citations
F, -- very rare -- Having only single citation in OLD or L+S
I, -- inscription -- Only citation is inscription
M, -- graffiti -- Presently not much used
N -- Pliny -- Things that appear only in Pliny Natural History
);

(Of course, Whitaker knows about Diederich's work -- he is the one who OCR'd Diederich's 1939 thesis and put it online.)

So, we're pleased to report that the Profile of Croatian Neo-Latin Project converted Whitaker's DICTPAGE.RAW to a MySQL table, and learned the following about how Whitaker's ten frequency categories are distributed among the 39,225 lemmata in his wordlist:

  1. X (Unknown or unspecified): 0

  2. A (very freq): 2134

  3. B (frequent): 2747

  4. C (common): 5113

  5. D (lesser): 8365

  6. E (uncommon): 11193

  7. F (very rare): 7974

  8. I (inscription): 430

  9. M (graffiti): 0

  10. N (Pliny): 1269

  11. Total: 39225


Now we have something to compare. It is interesting to note that most words are uncommon.

[Further reading.] There is a recent publication, Joseph Denooz, Nouveau lexique fréquentiel de latin. Alpha-Omega. Reihe A Bd 258. Hildesheim/Zürich/New York: Georg Olms Verlag, 2010. Pp. ix, 453. ISBN 9783487144733. €148.00. (reviewed recently on BMCR, with a crucial question: "A dictionary such as this is a tool: so what can this one be used for?").

[1] A sad update. Thinking about possible reasons for William Whitaker's absence from the internet, I consulted the obituaries, and found the following:

Colonel William A. Whitaker (USAF-Retired) passed away on Tuesday, December 14, 2010. While at DARPA, he worked on the computer language ADA. In retirement, he created the Latin-English translation software program, "Whitaker Words". (...)
Published in Midland Reporter-Telegram on December 21, 2010
Source here.


Τάνδε κατ' εὔδενδρον στείβων δρίος εἴρυσα χειρὶ
πτώσσουσαν βρομίας οἰνάδος ἐν πετάλοις,
ὄφρα μοι εὐερκεῖ καναχὰν δόμῳ ἔνδοθι θείη,
τερπνὰ δι' ἀγλώσσου φθεγγομένα στόματος.

Requiescat in pace.

Quantification

This morning we had to compile some numbers on the Croatiae auctores Latini collection. Here they are (also on the CroALa developer's blog):

  • 143 TEI XML files (including, alas, some duplicates)

  • 437.218 words

  • 29.637.450 characters

  • 16.465,25 Textkarten

  • 1029 Druckbogen


Last two strange categories belong to German printing tradition, which was influential in Croatian printing industry; we translated these terms (Textkarte = kartica teksta, Druckbogen = tiskarski arak), and use them still in text accounting.

[Technical note.] Numbers were produced by Linux wc command (cf. recipe) on all XML files currently in CroALa, also available on its Sourceforge page. The Linux one-liner for calculating number of characters and words in multiple XML files was simple:

wc *.xml | awk '{print $3-$1}'