Skip to content

[#295] - Search mask, facets and autocomplete for the TEI text edition - #316

Open
sebhofmann wants to merge 3 commits into
milestones/17-tei-praesentationsschichtfrom
issues/295-tei-search-mask
Open

[#295] - Search mask, facets and autocomplete for the TEI text edition#316
sebhofmann wants to merge 3 commits into
milestones/17-tei-praesentationsschichtfrom
issues/295-tei-search-mask

Conversation

@sebhofmann

Copy link
Copy Markdown
Collaborator

Closes #295, works on #296

Facettierte Suche in der TEI-Textedition, dazu die Indexierung der TEI-Dateien und ein Autocomplete für die Felder, deren Werte man nicht auswendig tippt.

Indexierung ([#296])

  • CMOTEIFileIndexAccumulator liest das TEI einer Textedition und schreibt seinen Inhalt in Solr-Felder des Dateidokuments: Text, Liedtext, Apparat und alles, was der teiHeader über Stück, Dichter, Versmaß und Textzeugen hält.
  • schema.xml bekommt diese Felder, je eine gestemmte Kopie pro Projektsprache (tei.lang.de.* usw.) und eine unveränderte Kopie für Facetten und Dropdowns (tei.facet.*), beide per Wildcard, damit ein neues TEI-Feld das Schema nicht anfasst.
  • CMOTEIFileIndexStrategy hält die TEI-Dateien aus /update/extract heraus. Wird der Inhalt einer Datei gesendet, muss jedes Feld ihres Dokuments als literal.*-Parameter in der URL reisen, was der Text einer Edition über das Header-Limit des Servlet-Containers treibt: es wurde bisher gar nichts davon indexiert. Die Extraktion ist auch nicht nötig, der Accumulator liest das TEI selbst und stellt den Text strukturiert und pro Element bereit.

Die Datei entscheidet sich am Wurzelelement, gelesen per StAX ohne DTD und externe Entities. Die Editions-TEIs sind UTF-16, die Kodierung kommt aus der Byte Order Mark.

Suchmaske und Facetten ([#295])

  • Neue Maske "Textedition" in der Editions-Suche. Gesucht werden die TEI-Felder über einen Join vom Dateidokument zur Textedition: Volltext, Volltext mit Apparat, Liedtext, Incipit, Art der Lesart, Vezin, Metrische Lizenz, Textdichter, Komponist, Lebensdaten, unsichere Dichterzuschreibung, Textzeuge.
  • Die Sidebar dieser Suche bekommt Facetten für Makâm, Usûl und Musikgattung (stehen an der Expression, zwei Schritte über den Katalog) sowie Textform, Versmaß und Reimschema (stehen am TEI). Jede braucht einen Join, deshalb rechnet sie die JSON Facet API: die Buckets werden in der Domäne der gejointen Dokumente gezählt, eine geschachtelte Facette joint zurück auf die Treffer und liefert die Zahl, die ein Leser erwartet.
  • Jeder Filter der Suche reist genau einmal als eigener Parameter, die Facetten verweisen darauf als {"param": "ffN"}. Ein Local-Param-Verweis ({!v=$ff0}) darf hier nicht stehen, das ist eine textuelle Substitution, bei der ein Wert am ersten Leerzeichen endet.
  • Semantik wie im übrigen Projekt: jeder angeklickte Eintrag ist ein eigener fq, alle müssen erfüllt sein. Deshalb dünnen die Facetten sich gegenseitig aus.
  • Was die Sidebar als Facette anbietet, steht nicht mehr in der Maske. Diese Werte sind Kategorien und Codes, die man besser auswählt als tippt, und die Facette zeigt dazu die Trefferzahl.

Autocomplete ([#295])

Felder, deren Werte man nicht auswendig kennt, schlagen sie beim Tippen vor: die transkribierten Namen der Dichter und Komponisten, die Siglen der Quellen und die Art der Lesart.

  • Textfelder werden über ihren Analyzer getroffen, hafiz findet also Hâfız-ı Şîrâzî. Die Werte kommen aus dem Highlighting von Solr, das von einem mehrwertigen Feld nur die passenden Werte liefert und nicht die Mitbewohner desselben Dokuments. Gefragt werden Wort und Wortanfang, weil die Textfelder gestemmt sind, eine Präfix-Anfrage aber nicht analysiert wird.
  • String-Felder werden genommen, wie sie stehen, Groß- und Kleinschreibung ignoriert: eine Facette mit facet.contains behält nur die Werte, die das Getippte enthalten, also auch mitten im Wert (ti findet addition).
  • Ab 2 Zeichen, 250 ms Tippruhe, maximal 10 Einträge, Pfeiltasten, Enter, Escape, Klick. Nach der Übernahme läuft die Suche neu. Als combobox/listbox/option mit aria-expanded und aria-activedescendant verdrahtet.
  • Aktiv in Textedition (Lyriker, Komponist, Textzeuge, Art der Lesart), Expression komplex und Liedtexte (Komponist, Lyriker, jeweils eingeschränkt auf die Personen mit dieser Rolle), Quelle (Mitwirkende, Veröffentlichungsinformationen), Literaturangabe und Edition (Name, Verlag) sowie Person (Name).

Sonstiges

Ein paar Zeichen in messages_de.properties waren als einzelne Bytes in die ASCII-Datei geschrieben und wurden seitdem als Ersetzungszeichen angezeigt, die sind mitrepariert.

Getestet

  • mvn clean install grün, 47 Tests, davon 18 für den Accumulator und 7 für die Index-Strategy.
  • Facetten gegen den lokalen Index geprüft: 67 anklickbare Einträge über fünf Facetten, keine Abweichung zwischen angezeigter Zahl und Treffern nach dem Klick.
  • Autocomplete im Browser gegen den echten Index geprüft: hafizHâfız-ı Şîrâzî (Klick und Tastatur), tmkTMKlii, tiaddition, ozbek → vier Namensvarianten, dede als Komponist → Dede Efendi mit 504 Treffern, isk und istanb in der Quellenmaske. Keine Konsolenfehler.

CMOTEIFileIndexAccumulator reads the TEI of a text edition and states its
content in solr fields of the document of the file: the text of the edition,
the lyrics, the apparatus and everything the teiHeader holds about the piece,
its poets, its metre and its witnesses. The schema gets those fields, a
stemmed copy per project language and an untouched copy for facets and drop
downs.

CMOTEIFileIndexStrategy keeps those files out of the extracting update
handler of solr. A file whose content is sent has to carry every field of its
document as a literal parameter of the url, which the text of an edition
blows beyond the header limit of the servlet container, so nothing of it was
indexed at all. The extraction is not needed either, the accumulator reads
the TEI itself and states its text structured and per element.
…text edition

The edition search gets a mask of its own, "Textedition": it searches the
fields of the TEI over a join from the document of the file to the text
edition it belongs to, in the full text, the apparatus, the lyrics, the
incipit, the metrical licences, the poets and composers and the witnesses.

The sidebar of that search gets facets for makam, usul and music genre, which
stand at the expression two steps away over the catalogue, and for text form,
metre and rhyme scheme, which stand at the TEI. Each of them needs a join, so
they are computed by the json facet api: the buckets are counted in the domain
of the joined documents and a nested facet joins back to the hits to state the
number a reader expects. Every filter of the search travels once as its own
parameter, the facets point at it as {"param": "ffN"}, which keeps the url
short. What the sidebar offers as a facet is left out of the mask, those values
are categories and codes and are better picked than typed.

Text fields whose values nobody knows by heart suggest them while the user
types, the transcribed names of the poets and composers, the sigla of the
sources and the type of a reading. A text field is matched by its analyzer, so
a name of the corpus is found without its diacritics as well, and the values
are read out of the highlighting of solr, which returns only the matching ones
of a multi valued field. A string field is matched as it stands, the case
ignored, by a facet which only keeps the values containing what was typed.

Also repairs some characters of the german messages, which were saved into the
ASCII file as single bytes and were shown as replacement characters since.
The url of a single hit is built from the parameters of the search, and the
ones which only feed the counters of the facets belong to the sidebar: the
json facet, the switch of the classic facets and every filter the counters
point at. The last ones were still carried along, so the browse view of a hit
got a handful of ff0, ff1 parameters it has no use for.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant