Author: ctschroeder (page 4 of 5)

Server updates – part of site down tonight

We are making some updates to the document application at data.copticscriptorium.org tonight (13 December 2016) approximately 7:30-8:30 pm Pacific time/10:30-11:30 Eastern time.  The service may be down.

You can still query and access our corpora in the ANNIS database at https://corpling.uis.georgetown.edu/annis/scriptorium .  That service will not be affected.  Thanks!

NEH White Paper (Preservations and Access Grant) published

We at Coptic SCRIPTORIUM have been fortunate to have received three grants from the National Endowment for the Humanities for our work.   We cannot thank the NEH enough for its support.  So much of what we have done over the past 2+ years could not have happened without this funding.

We just completed a White Paper paper for a Foundations grant from the Humanities Collections and Reference Resources program in the Division of Preservation and Access.  The grant, “Coptic SCRIPTORIUM: Digitizing a Corpus for Interdisciplinary Research in Ancient Egyptian,” ran from May 2104 until now.

Our White Paper documents our work and especially the standards and practices we developed for digitizing a pilot Coptic corpus.

If you want to know more about what truly interdisciplinary DH work looks like, check it out.  We try to break down the complexities of creating a digital corpus for research in linguistics, history, religious studies, biblical studies, manuscript studies.  We’ve got data models, workflows, digitization standards, transcription guidelines, and more all laid out for you here.

There is so much more to do; this is a only start.  Thanks to everyone who has had faith in our work.

White Paper, NEH Grant PW-51672-14 (Preservation and Access): “Coptic SCRIPTORIUM: Digitizing a Corpus for Interdisciplinary Research in Ancient Egyptian” 29 August 2016

Online lexicon linked to our corpora!

We have a great announcement today.  Along with our German research partners as part of the KELLIA project, we are releasing an online Coptic lexicon linked to our corpora.

For over three years, the Berlin-Brandenberg Academy of Sciences has been working on a digital lexicon for Coptic.  Frank Feder began the work.  Frank Feder began creating it, encoding definitions for Coptic lemmas in three languages: English, French, and German. The final entries were completed by Maxim Kupreyev at the academy and Julien Delhez in Göttingen.  The base lexicon file is encoded in TEI-XML.  This summer Amir Zeldes and his student, Emma Manning, created a web interface.  We will release the source code soon as part of the KELLIA project.

It may still need some refinements and updates, but we think it is a useful achievement that will help anyone interested in Coptic.

Entries have definitions in French, German, and English.

You can use the lexicon as a standalone website.  For the pilot launch, it’s on the Georgetown server, but make no mistake, this is major research outcome for the BBAW.

We’ve also linked the dictionary to our texts in Coptic SCRIPTORIUM.  You can click on the ANNIS icon in the dictionary entry to search all corpora in Coptic SCRIPTORIUM for that word.

lexicon-to-ANNIS The link also goes in the other direction.  In the normalized visualization of our texts, you can click on a word and get taken to the entry for that word’s lemma in the dictionary.  You can do this in the normalized visualization in our web application for reading and accessing texts (pictured below), or in the normalized visualization embedded in the ANNIS tool.

Screen Shot 2016-07-28 at 10.22.39 AM

Of course there will be refinements and developments to come.  We would love to hear your feedback on what works, what could work better, and where you find glitches.

On a more personal note, when Amir and I first came up with the idea for the project, we dreamed of creating a Perseus Digital Library for Coptic.  This dictionary is a huge step forward.  And honestly, I myself had almost nothing to do with this piece of the project.  It’s an example of the importance and power of collaboration.

New feature + texts in our corpora: Apophthegmata, I See Your Eagerness

We are very excited to release new versions of two of our corpora in time for the Coptic Congress.  And keep reading to learn about a new feature on our website.

As usually, we provide a diplomatic transcription of the texts’ manuscripts, normalized text for ease of reading, and an analytic visualization with the normalized text and part of speech tags in our web application.  Plus you’ll see buttons to search the corpora in our database or download our digital files.

Apophthegmata Patrum

The Apophthegmata Patrum now contains 36 published Sayings.  New ones include

This release also marks the first contributions of our newest editor, Dr. Dana Lampe.  Dana earned her Ph.D. at the Catholic University of America is beginning a postdoc at Creighton in the fall.

I See Your Eagerness

We also are releasing a huge new chunk of Shenoute’s sermon, I See Your Eagerness.  These texts were transcribed and collated primarily by David Brakke (with some by Stephen Emmel).  We thank David for his  generous donation of his transcriptions to the project!  Senior Editor Rebecca Krawiec has digitized and annotated these transcriptions.

Please begin your read of I See Your Eagerness with the fragment from codex MONB.GL 9-10.   Or you can search it in our search & visualization tool ANNIS.

We now have over 9000 words of this text digitized and annotated!

New: “Next” & “Previous” Buttons on Document visualizations

We’ve got a new feature in our web application:  the “next” and “previous” buttons near the top of the text.

“Next” is the next document for this work; if there is a lacuna, you’ll be taken to the next extant witness we’ve digitized.  If there are multiple, parallel witnesses, you’ll be taken to the witness we’ve identified as the best or clearest witness (typically based on the amount of lacunae).

The same is true for the “Previous” button.

If you want to review the parallel witness(es), check out the metadatum field for each document called “witness.”  If a parallel witness exists, it will be listed; if we have digitized the witness, the URN for the witness will be listed.  You can enter the URN in the box at the top of our website to retrieve the document.

Coptic SCRIPTORIUM at the Coptic Congress

Much of the Coptic SCRIPTORIUM team is in Claremont this week for the Congress of the International Association of Coptic Studies.

We started out with a pre-conference, 2-day workshop with our KELLIA partners from Germany, where we worked on sharing data and technologies across digital Coptic projects.  Look here soon for an announcement about a really cool fruit of our labors.

Thursday there are two panels, and Friday there are two workshops.

Thursday 2-4 pm Coptic Digital Studies (Burkle 16)

David Brakke chair 

Prof. Dr. Caroline Schroeder, Coptic SCRIPTORIUM: A Digital Platform for Research in Coptic Language and Literature

Dr. Christine Luckritz Marquis, Reimagining the Apopthegmata Patrum in a Digital Culture

Prof. Amir Zeldes, A Quantitative Approach to Syntactic Alternations in Sahidic

Dr. Rebecca Krawiec, Charting Rhetorical Choices in Shenoute: Abraham our Father and I See Your Eagerness as case-studies

Thursday 4:30-6:30 Coptic Digital Humanities (Burkle 16)

Caroline T. Schroeder, Chair

Dr. Paul Dilley, Coptic Scriptorium beyond the Manuscript: Towards a Distant Reading of Coptic Texts

Mr. So Miyagawa and Dr. Marco Büchler, Computational Analysis of Text Reuse in Shenoute and Besa

Mr. Uwe Sikora, Text Encoding – Opportunities and Challenges

Ms. Eliese-Sophia Lincke, Optical Character Recogition (OCR) for Coptic. Testing Automated Digitization of Texts with OCRopy

 

Friday 11-12:30 Workshop on Coptic Fonts & Coptic Bible (AA)

Christian Askeland, Frank Feder

Friday 4:30-6 Digital Tools for Beginners (Workshop on Coptic SCRIPTORIUM)

Caroline T. Schroeder, Amir Zeldes, Rebecca S. Krawiec

New publication: Raiders of the Lost Corpus in DHQ

Amir Zeldes and Caroline T. Schroeder have recently published an article in Digital Humanities Quarterly about the need for digital tools and a digitized corpus for Coptic, and research questions that drive Coptic SCRIPTORIUM.

“Raiders of the Lost Corpus” is freely available on the DHQ website as part of a special issue on Digital Methods and Classical Studies edited by Neil Coffee and Neil W. Bernstein.  Schroeder presented an earlier version of this paper at the Digital Classics conference at the University at Buffalo in 2013.

 

Full, machine-annotated New Testament Corpus updated

We’ve updated and re-released our fully machine-annotated New Testament corpus.  sahidica.nt V2.1.0 contains the Sahidica NT text from Warren Wells Sahidica online NT, with the following features:

  • Annotated with our latest NLP tools (part of speech tagger 1.9, tokenizer 4.1.0, language tagger and lemmatizer include lexical entries from the Database and Dictionary of Greek Loanwords in Coptic (DDGLC))
  • Now contains the morph layer (annotating compound words and Coptic morphs such ⲣⲉϥ- ⲙⲛⲧ- ⲁⲧ-)
  • Visualizations for linguistic analysis

Please keep in mind that this fully machine-annotated corpus is more accurate than previous versions but will nonetheless contain more errors than a corpus manually corrected by a human.

Search and queries

For searches and queries using our ANNIS database to find specific terms, for this corpus we recommend searching the normalized words using regular expressions (to capture instances of the desired word that may still be embedded in a Coptic bound group, instances that our tokenizer may have missed):

Lemma searches are now also possible.  You may wish to search for the lemma using regular expressions, as well, in order to find lemmas of some compound words.  For example, the following search will find entries containing ⲥⲱⲧⲙ in the lemma:

The results include various forms of ⲥⲱⲧⲙ (including ⲥⲟⲧⲙ) lemmatized the lexical entry “ⲥⲱⲧⲙ“, compound words lemmatized to ⲥⲱⲧⲙ or to a lexical entry containing ⲥⲱⲧⲙ, and some bound groups containing the word form ⲥⲱⲧⲙ, which our tokenizer did not catch:

Frequency table of normalized words lemmatized to swtm or a lemma form containing swtm (May 2016 Sahidica corpus)

Frequency table of normalized words lemmatized to ⲥⲱⲧⲙ or a lemma form containing ⲥⲱⲧⲙ (May 2016 Sahidica corpus)

As you can see, most of the hits are accurate (e.g., ⲥⲟⲧⲙ, ⲁⲧⲥⲱⲧⲙ, ⲣⲁⲧⲥⲱⲧⲙ, ⲣⲉϥⲥⲱⲧⲙ); some of the Coptic bound groups did not tokenize properly (e.g., ⲉⲡⲥⲱⲧⲙ, ⲙⲁⲣⲟⲩⲥⲱⲧⲙ).  We expect accuracy to increase as we incorporate more texts into our corpora that have been machine annotated and then manually edited.

Reading by individual chapter

You can also read these documents and see the linguistic analysis visualizations at data.copticscriptorium.org/urn:cts:copticLit:nt.  The first documents you will see (Gospel of Mark, 1 Corinthians) are manually annotated.  Scroll down for “New Testament,” which is the full, machine-annotated Sahidica New Testament.  Click on “Chapter” to read each chapter as normalized Coptic (with English translation as a pop-up when you hover your cursor).  Click on “Analytic” for the normalized Coptic, part of speech analysis, and English translation for each chapter.  Please keep in mind the English translation provided is a free, open-access New Testament translation from the World English Bible; it is not a direct translation from the Coptic.

Note:  we know that our server is slow generating the documents for this corpus.  It may take several minutes to load; please be patient.  For faster access, use ANNIS.  Visualizations to read the chapters are available by clicking on the corpus and the icon for visualizations.

Accessing document visualizations of the Sahidica corpus via ANNIS

Accessing document visualizations of the Sahidica corpus via ANNIS

We hope this corpus is useful to researchers.

New Besa fragment published

We’ve published another small fragment of Besa on Coptic SCRIPTORIUM.  So Miyagawa has edited and translated the letter fragment known as On Lack of Food.  Read it online or search the letters of Besa we have published.

Annotation tools now include DDGLC Greek Loanword List

We are pleased to announce the release of our newest versions of some of our natural language processing tools for Coptic which incorporate the lemma list of loanwords developed by the Database and Dictionary of Greek Loanwords in Coptic (DDGLC).

The DDGLC is part of the KELLIA partnership between American and German digital Coptic projects funded by the NEH Office of Digital Humanities and the DFG.  The DDGLC, under the direction of Prof. Dr. Tonio Sebastian Richter, has been building a database of Greek loanwords in Coptic in order to facilitate the study of language contact, language borrowing, and multilingualism in Egypt.

We have integrated the Greek lemma list into our language of origin tagger, tokenizer and morphology analysis, and lemmatizer.

Our online natural language processing web service (which bundles together all of our NLP tools into one web application) also includes this new data from the DDGLC.

The Greek loanword list should greatly increase the accuracy of many of our tools.  If you use them, please let us know how it goes!

We at Coptic SCRIPTORIUM are grateful for this partnership and the generosity of the DDGLC team.

New born-digital edition of a Shenoute fragment

This winter we’ve released a new document we’ve been working on for a while.  It’s a born digital publication, in the sense that this document to our knowledge has never been published previously.  The edition and annotations here were produced by Elizabeth Platte (Reed College) and Rebecca S. Krawiec (Canisius College) directly from digital photographs of the manuscript for digital publication.

Read the manuscript transcription or the  normalized text, or query it in our database.

It’s a section of one of Shenoute’s texts for monks in volume three of his monastic Canons.  This 14-page (seven-folio) fragment now resides in the Bibliothèque Nationale in Paris and originally derives from the White Monastery codex known by the siglum MONB.YB.  We’ve released text and annotations for pages 307-320, which equate to the BN call number Ms Copte 130/2 ff. 51-57.  Digital photos are now available online at Gallica.

We’ve transcribed the text from images of the manuscript and then annotated it for manuscript information.  We’ve also broken the text down into the Coptic phrases known as “bound groups,” words, and morphs.  Then we’ve annotated it all for part of speech, loan words (Greek, Latin, etc.), and lemmas.

By “we” I mean primarily Platte and Krawiec .  Schroeder and Zeldes provided editorial review, as per our policy of having every published digital document reviewed by at least one editor.

As far as we know, this fragment has never been published; nor has any translation ever been published.  We don’t have a translation yet, either.

As the first born-digital edition, this document is an experiment for us.  Everything else we’ve worked with has been published in an edition, and sometimes even has an English translation that another scholar has published.  Even though we digitize from the original manuscript, previous editions and translations make the transcription, annotation, and editing process much easier.  This document is an unknown quantity.

This means we expect to have errors and welcome feedback on the document.

We also have no translation as of yet.  Our goal is to translate the document and then edit the transcription and annotations again as we work.  We hope to publish an essay on how the digital annotation process affected the creation of an edition.

In the meantime, use it to practice your Coptic.  Let us know if you find errors.  We’ll credit you.

Older posts Newer posts