Meta tags:
description= Genteelly Observing the Enemy since 2011;
Headings (most frequently used words):
the, python, what, exactly, does, for, do, step, in, of, historical, recent, skulking, holes, and, corners, world, siege, sale, cleaning, text, with, sabbatical, rear, view, mirror, you, might, be, millner, from, source, to, data, where, historians, are, 2017, have, mentioned, future, is, digital, early, modern, spain, on, budget, part, reminder, smh, 2019, computer, program, scholars, this, historian, posts, comments, categories, bibliography, archives, blogroll, parsing, unparseable, converting, semi, structured, document, into, files,
Text of the page (most frequently used words):
the (548), and (263), you (251), that (149), with (119), for (114), #python (74), but (74), can (70), have (68), from (59), this (58), your (53), all (51), code (50), into (45), text (44), more (44), are (44), data (43), what (39), like (38), some (38), just (35), those (35), will (31), want (31), which (30), one (29), list (28), other (27), use (26), not (26), then (26), there (25), out (25), they (25), create (24), about (24), each (24), history (23), information (23), digital (22), could (22), document (22), does (21), their (21), few (21), them (21), how (21), over (20), any (20), was (20), well (19), also (19), word (19), already (18), world (18), most (18), has (18), now (17), early (17), modern (17), may (17), spanish (17), been (17), programming (17), make (17), words (17), computer (17), need (16), would (16), lots (16), even (16), own (16), 2018 (15), war (15), siege (15), research (15), when (15), catalan (15), many (15), its (15), historians (15), specific (15), analysis (15), documents (15), maps (14), least (14), bit (14), had (14), entries (14), get (13), 2014 (13), might (13), see (13), through (13), because (13), time (13), lot (13), several (13), libraries (13), historical (12), 2012 (12), august (12), 2013 (12), 2015 (12), 2019 (12), two (12), using (12), another (12), people (12), much (12), things (12), way (12), learning (12), folio (12), online (11), book (11), exactly (11), next (11), our (11), find (11), these (11), long (11), after (11), here (11), should (11), analyze (11), same (11), letters (11), only (11), results (11), different (11), excel (11), december (10), january (10), 2016 (10), together (10), year (10), who (10), barcelona (10), were (10), where (10), every (10), used (10), done (10), his (10), know (10), above (10), probably (10), language (10), clean (10), source (10), don (10), work (10), etc (10), maybe (10), automate (10), web (10), september (9), 2017 (9), zotero (9), day (9), process (9), think (9), back (9), various (9), case (9), separate (9), than (9), since (9), years (9), map (9), important (9), set (9), first (9), between (9), mentioned (9), big (9), available (9), dictionary (9), files (9), item (9), convert (9), format (9), errors (9), july (8), spain (8), search (8), take (8), number (8), off (8), count (8), making (8), authors (8), databases (8), read (8), sure (8), allows (8), example (8), file (8), easily (8), string (8), means (8), date (8), require (8), page (8), website (7), comments (7), skulking (7), holes (7), corners (7), blog (7), march (7), june (7), october (7), sources (7), methodology (7), cleaning (7), new (7), fortress (7), once (7), figure (7), did (7), better (7), access (7), small (7), museum (7), before (7), say (7), course (7), yak (7), full (7), really (7), start (7), future (7), run (7), still (7), structured (7), look (7), journal (7), yet (7), common (7), person (7), numbers (7), comment (6), content (6), wars (6), november (6), february (6), note (6), library (6), great (6), sale (6), recent (6), someone (6), including (6), let (6), sense (6), following (6), very (6), spent (6), particular (6), particularly (6), thing (6), shaving (6), never (6), steps (6), keep (6), check (6), wikipedia (6), whatever (6), order (6), allow (6), change (6), easy (6), value (6), nlp (6), perform (6), changes (6), jupyter (6), notebook (6), notes (6), tools (6), created (6), numerous (6), type (6), project (6), without (6), low (6), ocr (6), program (6), nouns (6), versions (6), features (6), texts (6), based (6), haven (6), enough (6), programs (6), functions (6), started (5), site (5), com (5), name (5), view (5), free (5), bibliography (5), april (5), sieges (5), military (5), google (5), devonthink (5), sabbatical (5), posts (5), via (5), leave (5), around (5), possibly (5), learn (5), doesn (5), ways (5), madrid (5), imagine (5), three (5), end (5), include (5), getting (5), real (5), fact (5), across (5), hand (5), line (5), combine (5), details (5), graph (5), past (5), step (5), yourself (5), campaign (5), little (5), add (5), dates (5), structure (5), pretty (5), entry (5), dozen (5), good (5), parse (5), record (5), copy (5), result (5), automatically (5), bunch (5), place (5), individual (5), thousands (5), historian (5), hanging (5), hard (5), projects (5), places (5), output (5), question (5), machine (5), working (5), letter (5), manipulate (5), types (5), tcp (5), wordpress (4), write (4), 2011 (4), france (4), conference (4), millner (4), finding (4), held (4), part (4), worth (4), visit (4), especially (4), wouldn (4), otherwise (4), fortunately (4), beginning (4), events (4), citadel (4), tell (4), fort (4), 1640 (4), while (4), why (4), give (4), main (4), right (4), said (4), heritage (4), catalonia (4), similar (4), therefore (4), guess (4), instead (4), english (4), city (4), succession (4), whole (4), during (4), four (4), semi (4), names (4), something (4), actually (4), almost (4), come (4), simple (4), scraping (4), open (4), deal (4), basic (4), series (4), doing (4), months (4), key (4), sorts (4), resulting (4), loop (4), delimiter (4), multiple (4), cleaned (4), camp (4), info (4), transcription (4), others (4), looks (4), regular (4), expressions (4), parsing (4), precise (4), problem (4), terms (4), phrases (4), likely (4), put (4), github (4), fruit (4), generally (4), ocred (4), dirty (4), kind (4), form (4), proper (4), entities (4), interactive (4), workflow (4), along (4), mean (4), makes (4), science (4), rename (4), download (4), network (4), made (4), eebo (4), pages (4), index (4), processing (4), books (4), coding (4), scholars (4), 000 (4), email (3), subscribe (3), account (3), rohan (3), british (3), army (3), primary (3), officers (3), taking (3), marlborough (3), art (3), resources (3), professional (3), miscellaneous (3), interested (3), help (3), smh (3), panels (3), coming (3), topic (3), definitely (3), wife (3), less (3), force (3), decided (3), build (3), starting (3), dozens (3), 1705 (3), thought (3), always (3), correct (3), forward (3), difference (3), central (3), recently (3), referendum (3), final (3), performing (3), interesting (3), somebody (3), rest (3), languages (3), ended (3), duke (3), service (3), allied (3), point (3), fall (3), king (3), memory (3), him (3), close (3), aren (3), unlike (3), today (3), currently (3), salses (3), left (3), medieval (3), pain (3), scratch (3), later (3), century (3), again (3), mostly (3), flags (3), far (3), having (3), week (3), last (3), whether (3), sparql (3), battles (3), location (3), post (3), hopefully (3), play (3), try (3), knowledge (3), until (3), qgis (3), reading (3), often (3), eventually (3), reason (3), images (3), journals (3), lists (3), hit (3), bits (3), statistical (3), function (3), returns (3), split (3), methods (3), notice (3), delete (3), standard (3), won (3), editor (3), standardizing (3), regex (3), sample (3), preprocessing (3), nor (3), separated (3), adding (3), top (3), level (3), relational (3), assuming (3), eye (3), sophisticated (3), database (3), searching (3), compared (3), nice (3), days (3), paragraph (3), collection (3), publication (3), compare (3), quickly (3), task (3), textual (3), ones (3), examples (3), image (3), digitized (3), institutional (3), provides (3), corpus (3), variable (3), works (3), syntax (3), repeat (3), folder (3), questions (3), software (3), answer (3), curve (3), rules (3), semester (3), department (3), stats (3), manually (3), entire (3), given (3), both (3), area (3), pdf (3), table (3), websites (3), quantitative (3), edit (3), pre (3), complicated (3), feature (3), natural (3), published (3), humanities (3), behind (3), commands (3), print (3), worry (3), microsoft (3), version (3), packaged (3), vep (3), required (2), log (2), sign (2), subscribed (2), louis (2), smhblog (2), shot (2), lineages (2), archives (2), group (2), society (2), strategy (2), operations (2), graphics (2), fortifications (2), dead (2), conferences (2), battle (2), 18c (2), ememh (2), teaching (2), admin (2), categories (2), jostwald (2), rear (2), mirror (2), older (2), panelists (2), proposals (2), european (2), love (2), considered (2), tramping (2), olympics (2), ditch (2), taken (2), commemorating (2), random (2), overlooking (2), submission (2), reapers (2), revolt (2), centuries (2), human (2), rights (2), torture (2), surprise (2), castle (2), exhibition (2), montjuïc (2), interactivity (2), says (2), opening (2), maintained (2), fortresses (2), below (2), hill (2), park (2), either (2), named (2), fast (2), built (2), decades (2), chapter (2), provide (2), community (2), however (2), government (2), support (2), population (2), against (2), symbol (2), resistance (2), relationship (2), national (2), tourism (2), creating (2), overall (2), saw (2), southern (2), philippe (2), berwick (2), fighting (2), everywhere (2), else (2), besiege (2), rather (2), austrian (2), karl (2), mind (2), clearly (2), earlier (2), familiar (2), minute (2), went (2), tolédo (2), experience (2), display (2), castilian (2), franco (2), thirty (2), keeping (2), track (2), gateway (2), model (2), nationalism (2), away (2), golden (2), video (2), flag (2), prehistory (2), undoubtedly (2), life (2), sticks (2), moderns (2), section (2), john (2), mention (2), guessing (2), speculation (2), historiographical (2), experienced (2), virtual (2), gives (2), moving (2), down (2), impression (2), bread (2), infinitely (2), league (2), looking (2), culture (2), budget (2), europe (2), meantime (2), query (2), listed (2), wikidata (2), curious (2), skills (2), scientist (2), gotta (2), granular (2), https (2), sheets (2), speed (2), times (2), faithful (2), usable (2), dataset (2), spreadsheets (2), warfare (2), downloaded (2), period (2), consider (2), bonus (2), internet (2), goes (2), overview (2), continue (2), yarn (2), surprising (2), allowing (2), oyster (2), extracting (2), depending (2), deep (2), aha (2), keywording (2), cycle (2), parsed (2), hundreds (2), metadata (2), visualization (2), modules (2), become (2), nested (2), items (2), strip (2), strings (2), associated (2), call (2), object (2), single (2), empty (2), seem (2), error (2), programmer (2), focus (2), original (2), easier (2), bbedit (2), formatting (2), seems (2), stage (2), lucky (2), scanned (2), missing (2), changing (2), spelling (2), contractions (2), spaces (2), needed (2), specifically (2), extracted (2), layout (2), towards (2), tabs (2), provided (2), argue (2), subject (2), piece (2), discrete (2), keyword (2), diary (2), drawn (2), being (2), converting (2), explain (2), lack (2), got (2), extract (2), limited (2), focusing (2), offer (2), possible (2), secondary (2), frequency (2), coordinates (2), according (2), criteria (2), visualizations (2), duration (2), similarly (2), groups (2), itself (2), further (2), talk (2), class (2), students (2), explore (2), method (2), pattern (2), parameter (2), input (2), practically (2), package (2), desired (2), field (2), topics (2), looked (2), algorithms (2), schedule (2), collecting (2), colleague (2), student (2), requests (2), interfaces (2), played (2), though (2), classification (2), sound (2), discuss (2), identify (2), context (2), calculate (2), 1700 (2), insert (2), update (2), records (2), mass (2), fields (2), relationships (2), nodes (2), topological (2), properties (2), international (2), visual (2), bokeh (2), spreadsheet (2), pandas (2), plot (2), charts (2), readable (2), requires (2), hyphenated (2), mistake (2), capitalize (2), uses (2), entity (2), corrections (2), hundred (2), ago (2), high (2), originals (2), amount (2), punctuation (2), dealing (2), ideas (2), previous (2), request (2), api (2), brute (2), apis (2), application (2), downloading (2), linked (2), copying (2), paste (2), pay (2), indexed (2), adobe (2), indexing (2), import (2), importance (2), literally (2), giant (2), kinds (2), digitally (2), fantasy (2), custom (2), usually (2), scenes (2), beyond (2), general (2), specialized (2), domain (2), tutorials (2), plenty (2), thanks (2), certain (2), pass (2), block (2), principles (2), attention (2), matter (2), quote (2), perfect (2), algorithm (2), expertise (2), emphasize (2), purchase (2), environments (2), automated (2), control (2), larger (2), upon (2), increase (2), inputs (2), manipulations (2), wish (2), summary (2), nothing (2), lets (2), layer (2), browser (2), variety (2), asking (2), matrix (2), 18th (2), share (2), myself (2), learned (2), geographical (2), summer (2), discoveries (2), installed (2), transcribed (2), england (2), volume (2), countries (2), genteelly (2), observing (2), enemy (2), design, loading, collapse, bar, manage, subscriptions, reader, report, privacy, join, 144, subscribers, quatorze, pike, oderint, dum, probent, perspectives, turenne, anno, domini, 1672, blogroll, wss, weapons, politics, veterans, travel, tactics, religion, recruitment, publishing, prussia, professionalism, pow, ottomans, netherlands, navy, enlightenment, mercenaries, medical, xiv, logistics, laws, italy, italian, ireland, intelligence, gtd, captains, gis, fwr, firearms, finance, engineers, emconventions, crusades, cavalry, caturday, casualties, britain, bcw, austria, artillery, sizes, armor, architecture, 30yw, 9yw, 7yw, uncategorized, historiography, select, category, wayne, lee, pradana, friends, propose, contact, seeking, forum, deadline, panel, paper, putting, alma, mater, ohio, state, columbus, theme, soldiers, civilians, cauldron, although, papers, germane, annual, reminder, caught, hosted, 1992, cue, archer, holdover, team, butts, dry, montjuic, locations, measurements, establish, meter, 1792, plaque, slightly, depressing, facts, occupying, stone, garrison, intimidate, town, coincidence, castilians, broke, bombard, restless, barcelonans, civil, became, violations, imprisonment, execution, political, prisoners, ends, quotation, universal, declaration, article, freedom, seek, involved, 1697, moat, monsters, guarding, ravelin, nope, guérite, smell, urine, photos, talking, zoom, descriptions, underwent, street, splay, embrasures, outward, funnel, themselves, flower, beds, aerial, bearings, complex, screenshot, surrounded, cable, car, funicular, walk, interest, sites, pronounced, mont, jew, initial, jewish, burial, ground, outcropping, mediterranean, transformed, lighthouse, successive, formidable, ethnographic, autonomous, northeast, known, implicitly, recognised, 1978, constitution, disagreements, tax, situation, hostility, independence, soared, polarising, outcome, remains, uncertain, previously, ruling, party, rebranded, non, binding, symbolic, pressure, politically, sensitive, significance, strained, johannes, venetia, identity, observations, edited, catherine, palmer, jacqueline, tivers, routledge, abstract, descriptive, noticed, witness, hot, press, study, catalonians, ours, explanation, oppressive, neighbor, defense, peeved, ordered, commander, slaughter, rebellion, capitulation, 11th, privileges, revoked, downhill, digress, surrender, meant, french, loaned, exposing, third, decade, fourth, brief, unsuccessful, attack, bent, eluded, large, exhibit, catastrophe, 1714, largely, 1713, crazy, catalans, kept, candidate, carlos, iii, abandoned, imperial, throne, 1711, becoming, holy, roman, emperor, appears, former, digs, remain, refused, acknowledge, felipe, stayed, death, witnessed, viewing, crypt, vienna, tomb, celebrates, 1706, liberation, kapuzinergruft, iberian, western, front, portuguese, border, animated, holdings, iberia, narrative, forces, managed, repulse, occupations, recapture, territory, battlefield, victory, almansa, obligingly, facilitated, reconquest, choosing, abandon, allies, contingent, captured, brihuega, museo, del, ejército, museu, història, catalunya, relevant, irregular, miquelets, rose, taxes, governance, olivares, home, songs, conflict, anthem, els, segadors, reaper, languedoc, visited, excursion, pyrenean, trouble, pirates, fun, came, origin, legend, senyera, 897, grateful, dipped, fingers, mortally, wounded, blood, dragged, shield, thanking, requested, killers, reward, differently, apparently, dating, 13c, prefer, stories, pen, sword, parchment, quill, ink, indeed, austrians, tale, crimson, streaked, half, waterfront, randomly, selected, destination, covers, region, hear, present, mannequin, cavemen, jump, 19th, rooms, kingdoms, principalities, shake, proverbial, stick, catalonian, speaking, struck, predominance, anglo, bookstores, geoffrey, parker, lynch, elliot, keegan, translations, histories, heavy, seeming, silence, senior, native, imperialism, reality, smartphone, headphone, fated, modernist, casa, bottló, photo, vertigo, anyway, attractions, mandatory, tourist, dollars, visiting, gaudí, standing, narrow, tower, waiting, wanting, contemplate, sagrada, familia, cathedral, tourists, must, proximity, bake, bâtards, wasn, rioting, spanked, liverpool, champions, survived, ugly, american, brooklynite, incident, train, arrived, northern, capital, seemed, sort, disagreement, parts, country, clue, proudly, hung, apartment, balconies, outnumbering, spying, vestiges, conquistadors, visits, castile, second, trip, side, conquerors, conquered, decide, eastern, belligerent, locating, efficiency, oriented, mindset, wherein, collects, ancient, explores, durations, attributes, cool, stuff, accurate, www, gokhan, collate, chart, heart, desires, wants, bureaucracy, inject, entered, factual, approximation, refine, beta, complete, defined, preparing, appeal, skulkers, assist, quixotic, quest, robust, envision, posting, aspects, combats, queries, lights, consist, phrase, describe, alludes, backward, move, sweater, shaved, sweaters, colorful, analogy, reilly, minor, tweaks, collections, newhailes, deane, takes, sucker, oysters, beauty, camps, filename, nest, dictionaries, layers, discussed, basics, 8th, precisely, target, hits, straddling, divide, pair, congratulations, ish, return, pickle, binary, passes, massage, values, efficient, fewer, lines, beginner, dict, comprehension, chained, encoding, splitting, hurdles, neophyte, discovered, pieces, fit, nutshell, within, interact, mileage, vary, 61404, muddled, curly, brackets, odd, equal, loss, care, robustly, relying, kindness, beggars, choosers, happen, reproduction, partial, tells, expanding, bother, rid, extra, stripped, standardized, identified, consists, carriage, distinctive, colons, ultimately, organized, per, appropriate, minutes, aside, examining, swearing, unstructured, consistent, scheme, judicious, unique, proof, concept, kindly, lawrence, smith, extent, mimicking, fidelity, 20th, maastricht, london, confusing, seen, explicitly, reinforces, best, passwords, solution, hone, tagging, transcriptions, requiring, wading, dross, likelihood, distracted, appear, filter, tested, physically, components, organize, old, quite, notecards, documented, journey, modified, applescripts, rudimentary, explode, rescue, tag, archival, begin, isolated, silos, timespan, needles, haystack, unparseable, automates, took, assistance, progress, show, preliminary, sweet, juicy, tantalizingly, reach, impediments, playing, pdfs, meaningful, systematic, slowly, catching, thus, pockets, multi, collaborators, position, sights, lower, taste, fruits, acquired, five, writing, diving, suggest, scan, decent, notepad, wrangler, killer, app, reads, statistics, identifies, rare, unusually, extracts, nationality, age, mentions, tables, graphs, locate, theaters, analyzed, runs, includes, comparisons, assigned, screen, chaining, objects, logic, copacetic, classes, voyant, among, heat, settings, parameters, performed, unable, faithfully, grant, agencies, recipients, submit, replicability, intriguing, rerun, exploration, super, anywhere, turn, extensible, related, automating, ask, drudgery, counting, sorting, revising, whom, hire, gets, burgeoning, classifying, superseded, neural, nets, specified, seeing, syllabus, meeting, removing, holidays, chair, reports, enrollment, assessment, surveys, scheduler, faculty, timeslots, university, departmental, scheduling, requirements, meets, administrative, school, batch, flexibility, mac, arcgis, boring, tasks, fledged, scriptable, portraits, facebook, reconstructing, soundscapes, genealogical, royal, genealogy, 16m, subsets, hathitrust, publications, periods, comparing, cited, affiliations, sciences, publish, ties, citation, networks, bibliometrics, biggie, extraction, cluster, segments, tend, collocated, keywords, sentiment, fuzzy, spelled, prose, phrasings, grammatical, structures, overuse, nltk, spacy, textacy, gensim, embeddings, calendars, tuesday, written, quick, calculation, calendar, datetime, dateutil, convertdate, dateparser, arrow, direction, contents, enter, pyzotero, interactions, sqlite, mysql, sqlite3, believe, triplets, verb, diagram, edges, measure, hubs, spokes, networkx, multiples, projection, whichever, aka, geocode, spatial, teams, gazetteers, georeferenced, geopandas, cartopy, mapping, explored, timeline, load, matplotlib, seaborn, humans, computers, detailed, description, ryan, cordell, hands, late, occurrence, outwork, misocred, mariborough, due, endings, changed, audit, difficult, capitalizations, liked, capitalization, clues, identifying, recognition, ideally, needs, special, handling, critical, bottleneck, idiosyncratic, historically, variant, usage, widely, varying, genre, vocabulary, sometimes, derived, ocring, irregularly, accuracy, rates, 100, ideal, quality, scans, average, scholar, acquire, ecco, paid, cheap, foreign, labor, theirs, social, scientists, born, everything, lowercase, stripping, imperfectly, retain, atomize, tokens, suggestions, welcome, useful, venues, 000s, skill, ted, underwood, jtb, raven, seriously, illustrate
Text of the page (random words):
nes i e most of the coding has already been done by the libraries authors so you just plug and play with some standard commands your specific data the specific variables you create and the specific commands combined in the specific order you choose figure out how to get your information into the code i e the computer s memory in a specific format based on its structure is it a string a list a dictionary manipulate the resulting data maybe you convert the string to a list after replacing certain features and adding others pass it on to the next block of code that does something else now that it s a list you can loop through each item and count those that start with the letter q do more things to the list and the count maybe if the count exceeds a certain threshold you send that word to another list then pass any of those results on to the next block of code until you end up with what you want to give a simple real world example maybe you ve started with a long string of text like this paragraph here you tokenize it into a list of individual words deciding how you want to deal with punctuation and contractions then count the words according to the letter they start with then plot a histogram showing the frequency of each letter python is also attractive to scholars because it s free insert joke about poor professor here its costlessness and open source ethos have encouraged hundreds of people to create free general libraries focusing on particular types of data and particular types of analysis along with specialized domain libraries for astronomers for geographers for audiologists for linguists for stock market analysts there is also a massive number of tutorials available online every year there are a dozen py conferences held all over the world and several hundred of the presentations are available on youtube including numerous 3 hour tutorials for beginners you can also check out the programming historian website which has numerous examples in python there are numerous cautions with programming in our case use python 3 not 2 7 but there are lots of resources that discuss those plenty of ways to get started in other words a final benefit of particular importance for humanities types is python s ability to convert words into numbers usually behind the scenes and highlight patterns using various statistical properties of text such powerful text functions allow businesses to data mine tweets and online content business demand seems to have juiced computer science research leading to lots of advanced natural language processing nlp features on top of those driven by the older linguistic and literary interests of academics so if you have a lot of digitized text or images and you want to clean analyze them beyond just reading each document one by one or manually cycling through search results one hit at a time then python is worth a look what exactly does python do for this historian so here s a list of the python projects i ve been working on and those i will be working on in the future a few are completed a few have draft code a few have some ideas sketched out with snippets of code and a couple are still in the fantasy phase many use the standard functions of off the shelf libraries while others require a bit more custom coding but they all should be viable projects time will tell semi automate a book index find all okay maybe most of the proper nouns in a pdf document along with which pdf page each occurred on then combine them together into a back of the book index format if you don t want to pay 1000 to have your book professionally indexed you could use word s or adobe s indexing feature which requires you to go through every sentence and identify which terms will need to be indexed or you can get 85 of that with python s nlp natural language processing libraries or you can import in a list of people places events and it will find those for you as with all programs things will get complicated the more edge and corner cases you try to address do i need to include a he on the next page with the full name on the previous page do i just combine together all consecutive pages into 34 39 or do i need to judge the importance of the headword to each page s discussion tough questions but this code will at the least give you a basis from which to tweak and judging from recent indexes in books published by highly reputable academic presses nobody cares about the art of indexing anymore some are literally headwords with a giant undifferentiated list of several dozen page numbers separated only by commas some don t even provide any kinds of topics only proper nouns of course the index may well die as more works are consumed digitally web scraping automate downloading content from a website either the text on pages entries in a list images of battle paintings from wikipedia or linked files maybe automate the sparql query on historical battles i posted about awhile back or download a bunch of letters from a site that puts each letter on its own separate page maybe automate scraping publication abstracts from a website based off records in zotero with a library like beautifulsoup web form entry i d like to create code that would automate copying bib info from zotero author title date pages etc and then paste it into our library s online ill form in the respective webform fields which of course aren t in the same order as the zotero field order that means a bunch of cutting and pasting for every request look up associated information on an entity person place organization with linked open data e g find the date of birth for person x mentioned in your text document via the web rdflib added by request get it apis numerous institutional websites allow you to access their online data more directly through api rather than using brute force scraping to harvest information from their pages html those websites have apis application program interfaces to allow more sophisticated downloading of information you can search their site for api to see if they offer it requests convert information into data my two previous posts on aha department enrollments and parsing long notes illustrate how this can be done with python code that converts information into data is particularly important low hanging fruit for historians since we lack a lot of already digitized datasets this kind of code allows us to create them with our own data clean dirty ocr text correct ocr errors and generally make a document more readable by humans and computers a good detailed description is ryan cordell s q i jtb the raven taking dirty ocr seriously this requires a lot of hands on work with the code which i ve been doing of late e g find every occurrence of out work and convert it to outwork so we can count them all the same way find every misocred mariborough and convert it to marlborough there are big lists of common errors available to make this search and edit process a bit more precise but since you ll never guess every word that might be hyphenated due to line endings finding every hyphenated word like be siege and converting it back to besiege is easy enough with regular expressions you can even create a list of all the changes your code makes i e create a dictionary with each mistake and what it was changed to if you want to audit the process more difficult are the questions of capitalizations case especially when our 17 18c authors liked to capitalize lots of nouns yet modern nlp uses capitalization as one of its clues for identifying proper nouns named entity recognition ideally you d have this code as a series of functions so that you can run the various corrections across an entire folder of documents based on what that source needs then you could use more code to check for any other errors or documents that require special handling i d argue that this is currently the most critical area for digital history the bottleneck in fact so few historians have their own sources in clean full text yet it s also a very idiosyncratic thing to program based on historically variant word usage and widely varying source genre vocabulary as well as the sometimes random errors derived from ocring irregularly set type from a few hundred years ago it d be great if ocr accuracy rates were 100 but that ideal would seem to require having high quality scans of the originals which your average scholar does not have and will likely never acquire because we don t actually own the originals as a result cleaning historical ocred text is the area with the least amount of pre made code available big projects like eebo and ecco paid cheap foreign labor to type theirs by hand we should also note that lots of social scientists for example talk about preprocessing text but by that they mean standardizing the spelling of text that s been born digital making everything lowercase stripping out punctuation etc historians need a lot of pre preprocessing first because we are dealing with imperfectly ocred text and if we want to retain a cleaned text copy of the original and not just atomize the text string into a list of word tokens then it s even more complicated suggestions welcome ted underwood has provided some useful ideas in various venues dealing with big data 10 000s of texts but his cleaning code on github is a bit above my skill level quantitative analysis load your spreadsheet into the pandas library and run stats plot charts pandas matplotlib seaborn bokeh for interactivity make interactive visualizations create interactive websites with your data and so on bokeh create a visual timeline drawn from data extracted from a document possibly interactive haven t explored this yet but i will mapping quickly make maps including small multiples in whatever projection with whichever features you want to display look up coordinates aka geocode and calculate spatial topological relationships it s good to see that there are several international project teams working on historical gazetteers and a number of groups have georeferenced some early modern maps as well geopandas cartopy network analysis create a network graph diagram of entities and relationships between those entities nodes and edges and measure the topological properties of the network who are the hubs the spokes the nodes networkx relational database interactions with sqlite and mysql sqlite3 i believe you can do the same with graph databases triplets of subject verb object but i haven t really looked at those zotero update and manipulate your zotero records i can already read zotero data into python and hand it off to other libraries for analysis as well as go in the other direction e g mass update fields back in zotero but i d like to create code that will take a pdf page or two of a book s table of contents and enter a separate record for each chapter in the book into zotero automatically adding in all the other book info pyzotero dates automatically calculate duration between two dates convert between os and ns and other historical calendars look up last tuesday s date when mentioned in a letter written on july 7 1700 you could probably automatically insert that date into the source document if desired maybe even do a quick calculation to see how few days you have left on your sabbatical calendar datetime dateutil convertdate dateparser arrow textual analysis this is a biggie for historians create a corpus of texts in your area create a list of people places and events to use for extraction cluster works or segments together by the topic they discuss see how often different authors texts use particular words phrases identify which words tend to be collocated with which other words keywords in context sentiment analysis etc did i mention fuzzy searching finding words that are spelled similarly or finding words that are used in the same context as a given word maybe you want to analyze your own prose which words phrasings grammatical structures do you overuse nltk spacy textacy gensim word embeddings bibliometrics and historiographical analysis from a secondary source extract all the people publications places and time periods dates mentioned and graph map them before comparing them with other authors or analyze the sources cited in the bibliography authors and affiliations years of publication languages etc the sciences have a lot of this already because they mostly publish journals and they re in databases like web of science this also ties into network analysis especially if you want to look at citation networks analyze words phrases from the 16m book hathitrust collection there s a website for that but you can also download the data or subsets at least genealogy parse genealogical data and analyze would be interesting for royal lineages and some work has already been done on that sound analysis haven t played with these but some people are into reconstructing soundscapes and the like image classification and analysis group together all the portraits of person x etc haven t played with these though you have similar classification algorithms in facebook etc lots of full fledged programs are also python scriptable e g both arcgis and qgis have python interfaces which means you can automate many of the boring tasks you need to perform when making more sophisticated maps clean up your computer files batch rename copy delete convert etc with much more flexibility than mac os x s rename function automate lots of administrative school work create a syllabus class schedule that lists the day of week and date for each meeting during the semester removing any holidays or other days off for that specific semester i ll be department chair next year and there are lots of stats and reports on enrollment assessment that i d like to automate collecting the data from databases surveys and then analyze them without me having to manually repeat the entire process every time a computer science colleague will have a student working on a course scheduler next semester given a department s faculty requests the available timeslots and a few dozen university and departmental scheduling requirements come up with a schedule that meets all those criteria or at least the most important ones now we have to do this by hand with excel but still and it s a real pain machine learning ai python is also one of the main languages used for this new burgeoning field for historians that might mean classifying documents and topics but i haven t looked into it enough to think about how it could be used some of the above mentioned libraries might well be superseded by machine learning libraries in the future where things like neural nets figure out their own algorithms without rules being specified by the programmer i think we re already seeing a little bit of that with nlp and those are just a few of the things you can do with python so whatever data related project you can think of there s probably a way to do it in python it s not just automating the things that y...
|