In this seminar I will demonstrate a new corpus search engine called pando, and the related graphical interface called flexicorp. Although PML-TQ is a powerful search engine that can search through dependency trees, it does not scale up enough to allow (rapidly) searching larger corpora that have been automatically parsed with UD, such as for instance ParlaMint with around 1B tokens, and InterCorp, with 5.5B tokens. The new pando engine follows the same strategy as CWB and Manatee, and keeps pace with those for sequence-based queries, but can run dependency-based queries with similar speed. Additionally, it has various features not supported by CWB, like persistent token names, real support for multi-value attributes, treatment of regions as first-class citizens, the options to use nested, overlapping, and discontinuous regions, and additional frequency options such as relative frequencies and (dependency) collocations. Pando was written explicitly to allow searching using TEITOK and by default provides JSON output for easy use as an API. But it can also be used as a command-line tool, called from Python or R, and I will show a modified version of Kontext that can query pando as well. Finally I will demonstrate how flexicorp has visualization options for various of the pando and TEITOK features, such as comparisons between queries or subcorpora, geolocation distribution, parallel corpora, time-aligned audio and facsimile images.
*** The talk will be delivered in person (MFF UK, Malostranské nám. 25, 4th floor, room S1) and will be streamed via Zoom. For details how to join the Zoom meeting, please write to sevcikova et ufal.mff.cuni.cz ***