Meta tags:
Headings (most frequently used words):
cross, linguistic, data, formats, why, what, design, principles, technology, history, cldf, specification, about, contact, info,
Text of the page (most frequently used words):
data (34), the (32), and (14), linguistic (13), cldf (12), for (12), with (10), cross (10), #formats (8), this (8), format (6), tools (6), should (6), from (5), language (5), specification (5), based (5), model (5), using (5), tabular (5), are (5), also (5), simple (4), new (4), standard (4), linguistics (4), main (3), possible (3), csv (3), exchange (3), databases (3), that (3), can (3), clld (3), methods (3), pipeline (3), style (3), use (3), typical (3), historical (3), semantics (3), like (3), reuse (3), automated (3), ontology (3), types (3), software (3), forkel (2), support (2), project (2), comparison (2), download (2), design (2), start (2), evolve (2), provide (2), thus (2), workshop (2), very (2), developed (2), has (2), built (2), top (2), core (2), framework (2), web (2), while (2), have (2), been (2), some (2), toolbox (2), workflows (2), unix (2), workflow (2), text (2), allow (2), easy (2), cognate (2), judgements (2), lingpy (2), could (2), phylogenetic (2), analysis (2), which (2), transformation (2), concerned (2), here (2), being (2), well (2), metadata (2), not (2), structure (2), but (2), compatible (2), 2018 (2), see (2), examples (2), publications (2), html (2), home (2), robert_forkel, eva, mpg, robert, contact, info, initiative, glottobank, consortium, max, planck, institute, evolutionary, anthropology, erc, computer, assisted, about, view, github, tar, ball, zip, file, simplicity, was, goal, under, consideration, will, starting, out, stable, baseline, further, evolution, following, discussions, first, leipzig, focused, idea, second, within, shown, many, different, same, attempt, externalise, trend, towards, computational, analyse, large, scale, interest, standardizing, particular, focus, around, time, sfm, used, developments, area, diversity, research, motivated, push, set, history, comparisons, procede, footsteps, bioinformatics, pipelines, may, point, formalized, common, suitable, line, available, does, extensibility, automatic, functionality, extended, post, processing, via, processes, sets, trees, represented, newick, ete, phyltr, one, goals, useful, delineation, makes, really, commands, seems, qlc, since, w3c, virtue, dialect, ideally, suited, combined, specify, syntax, serialization, much, sil, multi, dictionary, formatter, adds, hierarchical, markers, scenario, structures, make, analyses, mdf, json, vocabulary, technology, requires, specifies, just, stored, rigid, course, cannot, immediately, independently, mechanisms, let, understood, syntactically, compatibility, existing, standards, practice, always, kept, mind, entities, referenced, languages, through, their, glottocode, done, rather, than, duplicating, information, names, encoded, utf, files, both, editable, hand, amenable, reading, writing, preferably, linguist, expected, correctly, principles, dictionaries, datasets, wals, features, wordlists, more, complex, lexical, including, any, typically, analysed, quantitative, made, accessible, such, what, once, established, these, dataformats, become, foundation, only, instruction, material, spirit, typology, carpentry, decouple, development, standardized, necessary, why, advancing, sharing, comparative, sci, 180205, doi, 1038, sdata, 205, article, describing, list, changes, changelog, released,
Text of the page (random words):
cldf cross linguistic data formats home specification ontology html publications examples home specification ontology html publications examples cross linguistic data formats cldf 1 3 has been released see the changelog for a list of changes see also this article describing cldf forkel r et al cross linguistic data formats advancing data sharing and reuse in comparative linguistics sci data 5 180205 doi 10 1038 sdata 2018 205 2018 why to allow exchange of cross linguistic data and decouple development of tools and methods from that of databases standardized data formats are necessary once established these dataformats could become a foundation not only for tools but also for instruction material in the spirit of data carpentry for historical linguistics and linguistic typology what the main types of cross linguistic data we are concerned with here are any tabular data which is typically analysed using quantitative automated methods or made accessible using software tools like the clld framework such as wordlists or more complex lexical data including e g cognate judgements structure datasets e g wals features simple dictionaries design principles data should be both editable by hand and amenable to reading and writing by software preferably software the typical linguist can be expected to use correctly data should be encoded as utf 8 text files if entities can be referenced e g languages through their glottocode this should be done rather than duplicating information like language names compatibility with existing tools standards and practice should always be kept in mind automated re use requires that the standard specifies not just the structure but also the semantics of the data stored thus the cldf specification should be as rigid as possible of course new types of data cannot be immediately compatible with independently developed tools so the cldf standard should also provide mechanisms to let data types evolve well understood semantics while being syntactically compatible from the start technology since we are concerned with tabular data here cldf is built on w3c s model for tabular data and metadata on the web and metadata vocabulary for tabular data this model by virtue of being a json ld dialect is ideally suited to be combined with an ontology to specify syntax as well as semantics of a data serialization format much like mdf sil s multi dictionary formatter adds a hierarchical data model on top of toolbox standard format markers to support a data reuse scenario cldf structures cross linguistic data to make automated reuse in typical analyses in historical linguistics possible one of the main goals of the cldf specification is a useful delineation of data and tools using a csv based format makes it really easy to use this data in a unix style pipeline of data transformation commands this pipeline style of data transformation and analysis seems to be at the core of typical workflows e g in historical linguistics e g lingpy or qlc if suitable text and line based formats are available this pipeline style does also allow for easy extensibility e g a workflow for automatic cognate judgements based on lingpy functionality could be extended with phylogenetic analysis and post processing via phyltr which processes sets of phylogenetic trees represented in the newick format or ete if cross linguistic comparisons procede in the footsteps of bioinformatics workflows based on unix pipelines may at some point be formalized using a common workflow language history while data formats to exchange linguistic data have been around for some time e g the sfm or standard format used by toolbox new developments in the area of language diversity research have motivated this push for a new set of formats a new interest in standardizing tabular data on the web with a particular focus on csv a trend towards using computational methods to analyse large scale cross linguistic data the clld framework developed within the clld project has shown that many different cross linguistic databases can be built on top of the same core data model cldf is an attempt to externalise this data model thus following up discussions from the first workshop on language comparison with linguistic databases a second workshop in leipzig focused on the idea of a very simple csv based format to exchange very simple cross linguistic data simplicity was the main design goal from the start so the formats under consideration will evolve starting out as simple as possible with cldf 1 0 we provide a stable baseline for further evolution cldf specification download zip file download tar ball view on github about cldf is an initiative by the glottobank consortium with support from the max planck institute for evolutionary anthropology and the erc project computer assisted language comparison contact info robert forkel robert_forkel eva mpg de
|