Meta tags:
description= A character-wise tokenizer for morphologically rich languages - amir-zeldes/RFTokenizer;
Headings (most frequently used words):
uh, oh, navigation, training, file, saved, searches, rftokenizer, files, footer, search, code, repositories, users, issues, pull, requests, provide, feedback, amir, zeldes, menu, use, to, filter, your, results, more, quickly, folders, and, latest, commit, history, repository, installation, introduction, performance, requirements, using, about, releases, 11, packages, contributors, languages, command, line, importing, as, module, configuration, lexicon, frequency, other, options, resources, license, stars, watchers, forks,
Text of the page (most frequently used words):
the (73), for (35), and (35), you (26), file (20), word (20), this (19), data (19), can (17), with (16), #rftokenizer (16), github (15), training (15), hebrew (15), using (14), code (12), bert (12), model (11), forms (11), see (11), security (10), use (10), lexicon (10), from (10), segmentation (10), that (9), tokenizer (9), which (9), all (9), not (8), languages (8), view (8), train (8), features (8), line (8), heb (8), your (8), python (8), coptic (8), reload (7), character (7), tokens (7), used (7), format (6), license (6), column (6), conf (6), characters (6), lang (6), files (6), perfect (6), precision (6), recall (6), score (6), spmrl (6), available (6), amir (6), zeldes (6), enterprise (6), navigation (5), other (5), there (5), please (5), morphologically (5), rich (5), test (5), set (5), separated (5), example (5), are (5), frequency (5), tab (5), per (5), but (5), צלם (5), will (5), based (5), must (5), txt (5), arabic (5), search (5), information (4), community (4), was (4), error (4), while (4), loading (4), page (4), new (4), readme (4), wise (4), default (4), run (4), list (4), dataset (4), entries (4), delimited (4), text (4), one (4), following (4), two (4), super (4), tags (4), tag (4), sub (4), pos (4), wordform (4), have (4), עשרות (4), אנשים (4), מגיעים (4), splits (4), below (4), should (4), settings (4), name (4), models (4), under (4), open (4), requirements (4), treebank (4), out (4), com (4), more (4), paper (4), solutions (4), actions (4), quality (4), issues (4), support (4), time (3), stars (3), resources (3), specify (3), prune (3), during (3), form (3), multiple (3), case (3), first (3), include (3), given (3), simple (3), same (3), order (3), pipe (3), each (3), also (3), vowels (3), positions (3), last (3), useful (3), currently (3), small (3), classification (3), classifier (3), input (3), directory (3), scriptorium (3), without (3), ud_arabic (3), padt (3), scores (3), ud_hebrew (3), derived (3), http (3), www (3), morphological (3), setup (3), commit (3), insights (3), pull (3), requests (3), signed (3), another (3), window (3), refresh (3), session (3), sign (3), saved (3), documentation (3), feedback (3), grade (3), copilot (3), platform (3), explore (3), perform (2), personal (2), manage (2), terms (2), footer (2), 2026 (2), releases (2), latest (2), repository (2), forks (2), want (2), modify (2), dev (2), hyperparameter (2), optimization (2), comma (2), feature (2), importances (2), items (2), options (2), שמח (2), possible (2), sum (2), item (2), numbers (2), lines (2), required (2), complex (2), them (2), segments (2), noun (2), cplxpp (2), including (2), any (2), segment (2), second (2), current (2), next (2), assumed (2), provides (2), supply (2), מתאילנד (2), תאילנד (2), לישראל (2), ישראל (2), segmentations (2), pos_classes (2), base_letters (2), unused (2), diacritics (2), allowed (2), regex_tok (2), mapping (2), may (2), big (2), language (2), like (2), add (2), oov (2), configuration (2), flair (2), then (2), option (2), produce (2), learn (2), note (2), compiled (2), called (2), json (2), sm3 (2), tokenize_rf (2), need (2), provide (2), rf_tokenize (2), string (2), import (2), my_tokenizer (2), read (2), tokenized (2), pip (2), trained (2), only (2), xgboost (2), prague (2), dependency (2), wiki5k (2), domain (2), predictions (2), https (2), universaldependencies (2), copticscriptorium (2), provided (2), dependencies (2), wikipedia (2), approach (2), proceedings (2), 15th (2), sigmorphon (2), workshop (2), computational (2), research (2), phonetics (2), phonology (2), morphology (2), 2018 (2), address (2), brussels (2), belgium (2), 101 (2), 110 (2), tool (2), very (2), globally (2), optimal (2), better (2), medium (2), install (2), installation (2), nlp (2), replication (2), results (2), commits (2), message (2), menu (2), projects (2), appearance (2), cancel (2), searches (2), repositories (2), business (2), advanced (2), developer (2), topics (2), source (2), customer (2), services (2), devops (2), app (2), merge (2), action, share, cookies, contact, docs, status, privacy, inc, lex, contributors, packages, jun, speed, report, watching, watchers, activity, about, different, classifiers, hyperparameters, cross, validation, routine, fixed, look, cross_val_test, ablate, certain, retraining, entire, after, evaluation, variable, outputted, how, much, withheld, simulate, unseen, deactivate, split, proportion, partition, 32304, 39546, שמט, 314, integer, pooling, sources, integers, taken, repeated, recommended, give, distinct, sequence, above, contains, possessor, clitic, therefore, pronoun, token, verb, צילם, assigned, third, lemmas, reserved, future, meaningless, previous, meaningful, preceding, context, corpus, trigrams, four, columns, shuffled, vbp, vbz, vbd, vbg, vbn, nnp, nns, nnps, classes, אבגדהוזחטיכלמנסעפצקרשתןםךףץ, אהוי, next_letter, המבלושכ, המבלשכ, המבלכ, והיךםן, הכנ, followed, boundary, positive, beginning, starting, negative, end, when, setting, allow, closed, vocabulary, affixes, regular, expressions, rule, tokenization, names, permanently, disable, collapsed, reduce, sparseness, especially, distinguishes, something, matres, lectionis, configure, here, optionally, consider, treated, rare, emoji, etc, these, attested, section, header, top, corresponding, brackets, some, usually, named, config, parser, property, wish, trains, conllu, its, placed, disjoint, else, over, reliance, always, right, since, magically, predicts, everything, correctly, has, already, seen, alternatively, fold, regime, important, seg, flair_pos_tagger, supplied, sm2, freqs, invoked, least, ideally, containing, categorized, expects, analyze, return, value, analyses, separator, ara, test_ara, encoding, utf8, print, once, installed, via, importing, module, separate, newline, instead, output, example_in, example_out, select, command, compatible, specific, hyperopt, pandas, numpy, scikit, needs, 952007602755999, 9797786292039166, 9637772194304858, 971712054042643, ud_coptic, 984703395100112, 974712865553518, 9873875635510364, 9810092767982903, 9907224634820371, 9851075565361279, 9845644983461963, 9848359525778881, 9821036106750393, 9761790182868142, 967103694874851, 9716201652496708, latter, 9918367346938776, 9885304659498207, 9864091559370529, 9874686716791979, clean, experimental, 9933281004709577, 9923298178331735, 9871244635193133, 9897202964379631, realistic, jointly, iahlt, performance, corpora, org, experiment, universal, version, made, earlier, 2014, shared, task, attribution, htb, inproceedings, author, title, characterwisewindowed, ebrew, booktitle, year, pages, characterwise, windowed, cite, refer, binary, predicted, border, relies, fast, accurate, little, resists, overfitting, represent, crf, layer, transition, lattices, similar, coherent, into, known, categories, leads, handling, amounts, 10k, 200k, examples, works, box, fairly, summer, 2019, either, highest, published, accuracy, official, part, ensemble, does, internal, such, space, words, contain, clitics, segmented, smaller, units, contained, introduction, clone, repo, pypi, pretrained, sahidic, bohairic, dialects, hebpipe, full, pipelines, mrls, history, date, folders, branches, master, additional, star, fork, change, notification, notifications, public, dismiss, alert, switched, accounts, resetting, focus, create, qualifiers, our, query, filter, quickly, submit, email, contacted, every, piece, take, seriously, syntax, tips, clear, users, jump, pricing, premium, ons, powered, collections, trending, archive, program, accelerator, maintainer, lab, programs, fund, developers, sponsors, partners, trust, center, forum, skills, ebooks, reports, events, webinars, stories, type, software, development, topic, industries, government, manufacturing, financial, healthcare, industry, cases, devsecops, modernization, nonprofits, startups, teams, enterprises, company, size, marketplace, changelog, blog, why, stop, leaks, before, they, start, secret, protection, secure, build, find, fix, vulnerabilities, application, enforce, changes, review, plan, track, work, instant, environments, codespaces, automate, workflow, workflows, integrate, external, tools, mcp, registry, direct, agents, issue, write, creation, toggle, skip, content,
Text of the page (random words):
github amir zeldes rftokenizer a character wise tokenizer for morphologically rich languages github skip to content navigation menu toggle navigation sign in appearance settings platform ai code creation github copilot write better code with ai github copilot app direct agents from issue to merge mcp registry new integrate external tools developer workflows actions automate any workflow codespaces instant dev environments issues plan and track work code review manage code changes code quality enforce quality at merge application security github advanced security find and fix vulnerabilities code security secure your code as you build secret protection stop leaks before they start explore why github documentation blog changelog marketplace view all features solutions by company size enterprises small and medium teams startups nonprofits by use case app modernization devsecops devops ci cd view all use cases by industry healthcare financial services manufacturing government view all industries view all solutions resources explore by topic ai software development devops security view all topics explore by type customer stories events webinars ebooks reports business insights github skills support services documentation customer support community forum trust center partners view all resources open source community github sponsors fund open source developers programs security lab maintainer community accelerator github stars archive program repositories topics trending collections enterprise enterprise solutions enterprise platform ai powered developer platform available add ons github advanced security enterprise grade security features copilot for business enterprise grade ai features premium support enterprise grade 24 7 support pricing search or jump to search code repositories users issues pull requests search clear search syntax tips provide feedback we read every piece of feedback and take your input very seriously include my email address so i can be contacted cancel submit feedback saved searches use saved searches to filter your results more quickly name query to see all available qualifiers see our documentation cancel create saved search sign in sign up appearance settings resetting focus you signed in with another tab or window reload to refresh your session you signed out in another tab or window reload to refresh your session you switched accounts on another tab or window reload to refresh your session dismiss alert message amir zeldes rftokenizer public notifications you must be signed in to change notification settings fork 8 star 32 code issues 0 pull requests 1 actions projects security and quality 0 insights additional navigation options code issues pull requests actions projects security and quality insights amir zeldes rftokenizer master branches tags go to file code open more actions menu folders and files name name last commit message last commit date latest commit history 66 commits 66 commits rftokenizer rftokenizer license md license md readme md readme md requirements txt requirements txt setup py setup py view all files repository files navigation readme license more items rftokenizer a character wise tokenizer for morphologically rich languages for replication of paper results see replication md for full nlp pipelines for morphologically rich languages mrls based on this tool see coptic http www github com copticscriptorium coptic nlp hebrew http www github com amir zeldes hebpipe pretrained models are provided for coptic sahidic and bohairic dialects arabic and hebrew installation rftokenizer is available for installation from pypi pip install rftokenizer or you can clone this repo and run python setup py install introduction this is a simple tokenizer for word internal segmentation in morphologically rich languages such as hebrew coptic or arabic which have big super tokens space delimited words which contain e g clitics that need to be segmented and sub tokens the smaller units contained in super tokens segmentation is based on character wise binary classification each character is predicted to have a following border or not the tokenizer relies on an xgboost classifier which is fast very accurate using little training data and resists overfitting solutions do not represent globally optimal segmentations there is no crf layer transition lattices or similar but at the same time a globally coherent segmentation of each string into known morphological categories is not required which leads to better oov item handling the tokenizer is optimal for medium amounts of data 10k 200k examples of word forms to segment and works out of the box with fairly simple dependencies and small model files see requirements for two languages as of summer 2019 rftokenizer either provides the highest published segmentation accuracy on the official test set hebrew or forms part of an ensemble which does so coptic to cite this tool please refer to the following paper zeldes amir 2018 a characterwise windowed approach to hebrew morphological segmentation in proceedings of the 15th sigmorphon workshop on computational research in phonetics phonology and morphology brussels belgium 101 110 inproceedings author amir zeldes title a characterwisewindowed approach to h ebrew morphological segmentation booktitle proceedings of the 15th sigmorphon workshop on computational research in phonetics phonology and morphology year 2018 address brussels belgium pages 101 110 the data provided for the hebrew segmentation experiment in this paper given in the data directory is derived from the universal dependencies version of the hebrew treebank which is made available under a cc by nc sa 4 0 license but using the earlier splits from the 2014 spmrl shared task for attribution information for the hebrew treebank see https github com universaldependencies ud_hebrew htb the out of domain wikipedia dataset from the paper called wiki5k and available in the data directory is available under the same terms as wikipedia coptic data is derived from coptic scriptorium corpora see more information at http www copticscriptorium org arabic data is derived from the prague arabic dependency treebank ud_arabic padt https github com universaldependencies ud_arabic padt performance realistic scores on the spmrl hebrew dataset ud_hebrew v1 splits using bert based predictions and lexicon data as features trained jointly on spmrl and other ud hebrew iahlt data perfect word forms 0 9933281004709577 precision 0 9923298178331735 recall 0 9871244635193133 f score 0 9897202964379631 clean experimental scores on the spmrl hebrew dataset ud_hebrew v1 splits using bert based predictions and lexicon data as features and training only on spmrl perfect word forms 0 9918367346938776 precision 0 9885304659498207 recall 0 9864091559370529 f score 0 9874686716791979 or the latter without bert perfect word forms 0 9821036106750393 precision 0 9761790182868142 recall 0 967103694874851 f score 0 9716201652496708 scores on hebrew wiki5k out of domain with bert train on spmrl perfect word forms 0 9907224634820371 precision 0 9851075565361279 recall 0 9845644983461963 f score 0 9848359525778881 prague arabic dependency treebank ud_arabic padt currently without bert perfect word forms 0 984703395100112 precision 0 974712865553518 recall 0 9873875635510364 f score 0 9810092767982903 coptic scriptorium ud_coptic scriptorium currently without bert perfect word forms 0 952007602755999 precision 0 9797786292039166 recall 0 9637772194304858 f score 0 971712054042643 requirements the tokenizer needs scikit learn numpy pandas xgboost flair only if bert is used and if you want to run hyperparameter optimization hyperopt compatible with python 2 or 3 but compiled models must be specific to python 2 3 can t use a model trained under python 2 with python 3 using command line to use the tokenizer include the model files e g heb sm3 and heb json in the tokenizer s directory or in models then select it using m heb and supply a text file to run segmentation on the input file should have one word form per line for segmentation python tokenize_rf py m heb example_in txt example_out txt input file format עשרות אנשים מגיעים מתאילנד לישראל output format עשרות אנשים מגיעים מ תאילנד ל ישראל you can also use the option n to separate segments using a newline instead of the pipe character importing as a module you can import rftokenizer once it is installed e g via pip for example from rftokenizer import rftokenizer my_tokenizer rftokenizer model ara data open test_ara txt encoding utf8 read tokenized my_tokenizer rf_tokenize data print tokenized note that rf_tokenize expects a list of word forms to analyze or a string with word forms separated by new lines the return value is a list of analyses separated by the separator default training to train a new model you will need at least a configuration file and a training file ideally you should also provide a lexicon file containing categorized sub tokens and super tokens and frequency information for sub tokens see below training is invoked like this python tokenize_rf py t m lang c conf l lexicon f freqs training this will produce lang sm3 and lang json the compiled model files or sm2 under python 2 if conf is not supplied it is assumed to be called lang conf if you wish to use bert features for classification you must first train a flair classifier using flair_pos_tagger py which trains on conllu data and name its model lang seg which should be placed in models then train rftokenizer using the bert option important note the data used to train the bert classifier must be disjoint from the data used to train rftokenizer or else it will produce over reliance rftokenizer will learn that bert is always right since bert magically predicts everything correctly given that it has already seen this training data alternatively you can use a k fold training regime configuration you must specify some settings for your model in a file usually named lang conf e g heb conf for hebrew this file is a config parser property file with the following format a section header at the top corresponding to your language model in brackets e g heb for hebrew base_letters characters to consider during classification all other characters are treated as _ useful for oov rare characters emoji s etc these characters should be attested in training optionally you may add vowels if the language distinguishes something like vowels including matres lectionis it can be useful to configure them here pos_classes a mapping of pos tags to collapsed pos tags in the lexicon in order to reduce sparseness especially if the tag set is big but training data is small see below for format unused comma separated list of feature names to permanently disable in this model diacritics not currently used regex_tok a set of regular expressions used for rule based tokenization e g for numbers see example below allowed mapping of characters that may be followed by a boundary at positive positions in the beginning of the word starting at 0 or negative positions at the end of the word 1 is the last character when this setting is used no other characters positions will allow splits useful for languages with a closed vocabulary of affixes see below for format example heb conf file for hebrew heb base_letters אבגדהוזחטיכלמנסעפצקרשתןםךףץ vowels אהוי unused next_letter diacritics ּ allowed 0 המבלושכ 1 המבלשכ 2 המבלכ 3 ה 1 והיךםן 2 הכנ regex_tok 0 9 a za z 1 ב ל מ כ ה 0 9 a za z 1 2 3 ב ל מ כ ה 0 9 a za z 1 2 if using pos classes pos_classes v vbp vbz vb vbd vbg vbn md n nn nnp ns nns nnps training file a two column text file with word forms in one column and pipe delimited segmentations in the second column עשרות עשרות אנשים אנשים מגיעים מגיעים מתאילנד מ תאילנד לישראל ל ישראל it is assumed that line order is meaningful i e each line provides preceding context for the next line if you have a shuffled corpus of trigrams you can also supply a four column training file with the columns previous wordform next wordform current wordform current wordform segmentation pipe separated in this case line order is meaningless lexicon file the lexicon file is a tab delimited text file with one word form per line and the pos tag assigned to that word in a second column a third column with lemmas is reserved for future use multiple entries per word are possible e g צלם noun צלם צלם verb צילם צלם cplxpp צל it is recommended but not required to include entries for complex super tokens and give them distinct tags e g the sequence צלם above contains two segments a noun and a possessor clitic it is therefore given the tag cplxpp complex including a personal pronoun this tag is not used for any simple sub token segment in the same lexicon frequency file the frequency file is a tab delimited text file with one word form per line and the frequency of that word as an integer multiple entries per word are possible if pooling data from multiple sources in which case the sum of integers is taken in the following example the frequency of the repeated first item is the sum of the numbers in the first two lines שמח 32304 שמח 39546 שמט 314 other training options you can specify a train test split proportion using e g p 0 2 default test partition is 0 1 of the data you can specify how much of the lexicon to prune during training for example with prune 0 2 by default 0 1 of the lexicon entries are withheld during training to simulate unseen items at test time you can set prune 0 0 to deactivate this variable importances can be outputted using i you can perform retraining on the entire dataset after evaluation of feature importances using r you can ablate certain features using a and a comma separated list of features hyperparameter optimization can be run with o if you want to test different classifiers modify default hyperparameters you can modify the cross validation code in the train routine or use a fixed dev set look for cross_val_test about a character wise tokenizer for morphologically rich languages resources readme license view license uh oh there was an error while loading please reload this page activity stars 32 stars watchers 3 watching forks 8 forks report repository releases 11 v3 0 0 new model format speed latest jun 15 2026 10 releases packages 0 uh oh there was an error while loading please reload this page uh oh there was an error while loading please reload this page contributors uh oh there was an error while loading please reload this page languages lex 99 8 other 0 2 footer 2026 github inc footer navigation terms privacy security status community docs contact manage cookies do not share my personal information you can t perform that action at this time
|