Meta tags:
description= Genteelly Observing the Enemy since 2011;
Headings (most frequently used words):
the, python, what, exactly, does, for, do, step, in, of, historical, recent, skulking, holes, and, corners, world, siege, sale, cleaning, text, with, sabbatical, rear, view, mirror, you, might, be, millner, from, source, to, data, where, historians, are, 2017, have, mentioned, future, is, digital, early, modern, spain, on, budget, part, reminder, smh, 2019, computer, program, scholars, this, historian, posts, comments, categories, bibliography, archives, blogroll, parsing, unparseable, converting, semi, structured, document, into, files,
Text of the page (most frequently used words):
the (548), and (263), you (251), that (149), with (119), for (114), #python (74), but (74), can (70), have (68), from (59), this (58), your (53), all (51), code (50), into (45), text (44), more (44), are (44), data (43), what (39), like (38), some (38), just (35), those (35), will (31), want (31), which (30), one (29), list (28), other (27), use (26), not (26), then (26), there (25), out (25), they (25), create (24), about (24), each (24), history (23), information (23), digital (22), could (22), document (22), does (21), their (21), few (21), them (21), how (21), over (20), any (20), was (20), well (19), also (19), word (19), already (18), world (18), most (18), has (18), now (17), early (17), modern (17), may (17), spanish (17), been (17), programming (17), make (17), words (17), computer (17), need (16), would (16), lots (16), even (16), own (16), 2018 (15), war (15), siege (15), research (15), when (15), catalan (15), many (15), its (15), historians (15), specific (15), analysis (15), documents (15), maps (14), least (14), bit (14), had (14), entries (14), get (13), 2014 (13), might (13), see (13), through (13), because (13), time (13), lot (13), several (13), libraries (13), historical (12), 2012 (12), august (12), 2013 (12), 2015 (12), 2019 (12), two (12), using (12), another (12), people (12), much (12), things (12), way (12), learning (12), folio (12), online (11), book (11), exactly (11), next (11), our (11), find (11), these (11), long (11), after (11), here (11), should (11), analyze (11), same (11), letters (11), only (11), results (11), different (11), excel (11), december (10), january (10), 2016 (10), together (10), year (10), who (10), barcelona (10), were (10), where (10), every (10), used (10), done (10), his (10), know (10), above (10), probably (10), language (10), clean (10), source (10), don (10), work (10), etc (10), maybe (10), automate (10), web (10), september (9), 2017 (9), zotero (9), day (9), process (9), think (9), back (9), various (9), case (9), separate (9), than (9), since (9), years (9), map (9), important (9), set (9), first (9), between (9), mentioned (9), big (9), available (9), dictionary (9), files (9), item (9), convert (9), format (9), errors (9), july (8), spain (8), search (8), take (8), number (8), off (8), count (8), making (8), authors (8), databases (8), read (8), sure (8), allows (8), example (8), file (8), easily (8), string (8), means (8), date (8), require (8), page (8), website (7), comments (7), skulking (7), holes (7), corners (7), blog (7), march (7), june (7), october (7), sources (7), methodology (7), cleaning (7), new (7), fortress (7), once (7), figure (7), did (7), better (7), access (7), small (7), museum (7), before (7), say (7), course (7), yak (7), full (7), really (7), start (7), future (7), run (7), still (7), structured (7), look (7), journal (7), yet (7), common (7), person (7), numbers (7), comment (6), content (6), wars (6), november (6), february (6), note (6), library (6), great (6), sale (6), recent (6), someone (6), including (6), let (6), sense (6), following (6), very (6), spent (6), particular (6), particularly (6), thing (6), shaving (6), never (6), steps (6), keep (6), check (6), wikipedia (6), whatever (6), order (6), allow (6), change (6), easy (6), value (6), nlp (6), perform (6), changes (6), jupyter (6), notebook (6), notes (6), tools (6), created (6), numerous (6), type (6), project (6), without (6), low (6), ocr (6), program (6), nouns (6), versions (6), features (6), texts (6), based (6), haven (6), enough (6), programs (6), functions (6), started (5), site (5), com (5), name (5), view (5), free (5), bibliography (5), april (5), sieges (5), military (5), google (5), devonthink (5), sabbatical (5), posts (5), via (5), leave (5), around (5), possibly (5), learn (5), doesn (5), ways (5), madrid (5), imagine (5), three (5), end (5), include (5), getting (5), real (5), fact (5), across (5), hand (5), line (5), combine (5), details (5), graph (5), past (5), step (5), yourself (5), campaign (5), little (5), add (5), dates (5), structure (5), pretty (5), entry (5), dozen (5), good (5), parse (5), record (5), copy (5), result (5), automatically (5), bunch (5), place (5), individual (5), thousands (5), historian (5), hanging (5), hard (5), projects (5), places (5), output (5), question (5), machine (5), working (5), letter (5), manipulate (5), types (5), tcp (5), wordpress (4), write (4), 2011 (4), france (4), conference (4), millner (4), finding (4), held (4), part (4), worth (4), visit (4), especially (4), wouldn (4), otherwise (4), fortunately (4), beginning (4), events (4), citadel (4), tell (4), fort (4), 1640 (4), while (4), why (4), give (4), main (4), right (4), said (4), heritage (4), catalonia (4), similar (4), therefore (4), guess (4), instead (4), english (4), city (4), succession (4), whole (4), during (4), four (4), semi (4), names (4), something (4), actually (4), almost (4), come (4), simple (4), scraping (4), open (4), deal (4), basic (4), series (4), doing (4), months (4), key (4), sorts (4), resulting (4), loop (4), delimiter (4), multiple (4), cleaned (4), camp (4), info (4), transcription (4), others (4), looks (4), regular (4), expressions (4), parsing (4), precise (4), problem (4), terms (4), phrases (4), likely (4), put (4), github (4), fruit (4), generally (4), ocred (4), dirty (4), kind (4), form (4), proper (4), entities (4), interactive (4), workflow (4), along (4), mean (4), makes (4), science (4), rename (4), download (4), network (4), made (4), eebo (4), pages (4), index (4), processing (4), books (4), coding (4), scholars (4), 000 (4), email (3), subscribe (3), account (3), rohan (3), british (3), army (3), primary (3), officers (3), taking (3), marlborough (3), art (3), resources (3), professional (3), miscellaneous (3), interested (3), help (3), smh (3), panels (3), coming (3), topic (3), definitely (3), wife (3), less (3), force (3), decided (3), build (3), starting (3), dozens (3), 1705 (3), thought (3), always (3), correct (3), forward (3), difference (3), central (3), recently (3), referendum (3), final (3), performing (3), interesting (3), somebody (3), rest (3), languages (3), ended (3), duke (3), service (3), allied (3), point (3), fall (3), king (3), memory (3), him (3), close (3), aren (3), unlike (3), today (3), currently (3), salses (3), left (3), medieval (3), pain (3), scratch (3), later (3), century (3), again (3), mostly (3), flags (3), far (3), having (3), week (3), last (3), whether (3), sparql (3), battles (3), location (3), post (3), hopefully (3), play (3), try (3), knowledge (3), until (3), qgis (3), reading (3), often (3), eventually (3), reason (3), images (3), journals (3), lists (3), hit (3), bits (3), statistical (3), function (3), returns (3), split (3), methods (3), notice (3), delete (3), standard (3), won (3), editor (3), standardizing (3), regex (3), sample (3), preprocessing (3), nor (3), separated (3), adding (3), top (3), level (3), relational (3), assuming (3), eye (3), sophisticated (3), database (3), searching (3), compared (3), nice (3), days (3), paragraph (3), collection (3), publication (3), compare (3), quickly (3), task (3), textual (3), ones (3), examples (3), image (3), digitized (3), institutional (3), provides (3), corpus (3), variable (3), works (3), syntax (3), repeat (3), folder (3), questions (3), software (3), answer (3), curve (3), rules (3), semester (3), department (3), stats (3), manually (3), entire (3), given (3), both (3), area (3), pdf (3), table (3), websites (3), quantitative (3), edit (3), pre (3), complicated (3), feature (3), natural (3), published (3), humanities (3), behind (3), commands (3), print (3), worry (3), microsoft (3), version (3), packaged (3), vep (3), required (2), log (2), sign (2), subscribed (2), louis (2), smhblog (2), shot (2), lineages (2), archives (2), group (2), society (2), strategy (2), operations (2), graphics (2), fortifications (2), dead (2), conferences (2), battle (2), 18c (2), ememh (2), teaching (2), admin (2), categories (2), jostwald (2), rear (2), mirror (2), older (2), panelists (2), proposals (2), european (2), love (2), considered (2), tramping (2), olympics (2), ditch (2), taken (2), commemorating (2), random (2), overlooking (2), submission (2), reapers (2), revolt (2), centuries (2), human (2), rights (2), torture (2), surprise (2), castle (2), exhibition (2), montjuïc (2), interactivity (2), says (2), opening (2), maintained (2), fortresses (2), below (2), hill (2), park (2), either (2), named (2), fast (2), built (2), decades (2), chapter (2), provide (2), community (2), however (2), government (2), support (2), population (2), against (2), symbol (2), resistance (2), relationship (2), national (2), tourism (2), creating (2), overall (2), saw (2), southern (2), philippe (2), berwick (2), fighting (2), everywhere (2), else (2), besiege (2), rather (2), austrian (2), karl (2), mind (2), clearly (2), earlier (2), familiar (2), minute (2), went (2), tolédo (2), experience (2), display (2), castilian (2), franco (2), thirty (2), keeping (2), track (2), gateway (2), model (2), nationalism (2), away (2), golden (2), video (2), flag (2), prehistory (2), undoubtedly (2), life (2), sticks (2), moderns (2), section (2), john (2), mention (2), guessing (2), speculation (2), historiographical (2), experienced (2), virtual (2), gives (2), moving (2), down (2), impression (2), bread (2), infinitely (2), league (2), looking (2), culture (2), budget (2), europe (2), meantime (2), query (2), listed (2), wikidata (2), curious (2), skills (2), scientist (2), gotta (2), granular (2), https (2), sheets (2), speed (2), times (2), faithful (2), usable (2), dataset (2), spreadsheets (2), warfare (2), downloaded (2), period (2), consider (2), bonus (2), internet (2), goes (2), overview (2), continue (2), yarn (2), surprising (2), allowing (2), oyster (2), extracting (2), depending (2), deep (2), aha (2), keywording (2), cycle (2), parsed (2), hundreds (2), metadata (2), visualization (2), modules (2), become (2), nested (2), items (2), strip (2), strings (2), associated (2), call (2), object (2), single (2), empty (2), seem (2), error (2), programmer (2), focus (2), original (2), easier (2), bbedit (2), formatting (2), seems (2), stage (2), lucky (2), scanned (2), missing (2), changing (2), spelling (2), contractions (2), spaces (2), needed (2), specifically (2), extracted (2), layout (2), towards (2), tabs (2), provided (2), argue (2), subject (2), piece (2), discrete (2), keyword (2), diary (2), drawn (2), being (2), converting (2), explain (2), lack (2), got (2), extract (2), limited (2), focusing (2), offer (2), possible (2), secondary (2), frequency (2), coordinates (2), according (2), criteria (2), visualizations (2), duration (2), similarly (2), groups (2), itself (2), further (2), talk (2), class (2), students (2), explore (2), method (2), pattern (2), parameter (2), input (2), practically (2), package (2), desired (2), field (2), topics (2), looked (2), algorithms (2), schedule (2), collecting (2), colleague (2), student (2), requests (2), interfaces (2), played (2), though (2), classification (2), sound (2), discuss (2), identify (2), context (2), calculate (2), 1700 (2), insert (2), update (2), records (2), mass (2), fields (2), relationships (2), nodes (2), topological (2), properties (2), international (2), visual (2), bokeh (2), spreadsheet (2), pandas (2), plot (2), charts (2), readable (2), requires (2), hyphenated (2), mistake (2), capitalize (2), uses (2), entity (2), corrections (2), hundred (2), ago (2), high (2), originals (2), amount (2), punctuation (2), dealing (2), ideas (2), previous (2), request (2), api (2), brute (2), apis (2), application (2), downloading (2), linked (2), copying (2), paste (2), pay (2), indexed (2), adobe (2), indexing (2), import (2), importance (2), literally (2), giant (2), kinds (2), digitally (2), fantasy (2), custom (2), usually (2), scenes (2), beyond (2), general (2), specialized (2), domain (2), tutorials (2), plenty (2), thanks (2), certain (2), pass (2), block (2), principles (2), attention (2), matter (2), quote (2), perfect (2), algorithm (2), expertise (2), emphasize (2), purchase (2), environments (2), automated (2), control (2), larger (2), upon (2), increase (2), inputs (2), manipulations (2), wish (2), summary (2), nothing (2), lets (2), layer (2), browser (2), variety (2), asking (2), matrix (2), 18th (2), share (2), myself (2), learned (2), geographical (2), summer (2), discoveries (2), installed (2), transcribed (2), england (2), volume (2), countries (2), genteelly (2), observing (2), enemy (2), design, loading, collapse, bar, manage, subscriptions, reader, report, privacy, join, 144, subscribers, quatorze, pike, oderint, dum, probent, perspectives, turenne, anno, domini, 1672, blogroll, wss, weapons, politics, veterans, travel, tactics, religion, recruitment, publishing, prussia, professionalism, pow, ottomans, netherlands, navy, enlightenment, mercenaries, medical, xiv, logistics, laws, italy, italian, ireland, intelligence, gtd, captains, gis, fwr, firearms, finance, engineers, emconventions, crusades, cavalry, caturday, casualties, britain, bcw, austria, artillery, sizes, armor, architecture, 30yw, 9yw, 7yw, uncategorized, historiography, select, category, wayne, lee, pradana, friends, propose, contact, seeking, forum, deadline, panel, paper, putting, alma, mater, ohio, state, columbus, theme, soldiers, civilians, cauldron, although, papers, germane, annual, reminder, caught, hosted, 1992, cue, archer, holdover, team, butts, dry, montjuic, locations, measurements, establish, meter, 1792, plaque, slightly, depressing, facts, occupying, stone, garrison, intimidate, town, coincidence, castilians, broke, bombard, restless, barcelonans, civil, became, violations, imprisonment, execution, political, prisoners, ends, quotation, universal, declaration, article, freedom, seek, involved, 1697, moat, monsters, guarding, ravelin, nope, guérite, smell, urine, photos, talking, zoom, descriptions, underwent, street, splay, embrasures, outward, funnel, themselves, flower, beds, aerial, bearings, complex, screenshot, surrounded, cable, car, funicular, walk, interest, sites, pronounced, mont, jew, initial, jewish, burial, ground, outcropping, mediterranean, transformed, lighthouse, successive, formidable, ethnographic, autonomous, northeast, known, implicitly, recognised, 1978, constitution, disagreements, tax, situation, hostility, independence, soared, polarising, outcome, remains, uncertain, previously, ruling, party, rebranded, non, binding, symbolic, pressure, politically, sensitive, significance, strained, johannes, venetia, identity, observations, edited, catherine, palmer, jacqueline, tivers, routledge, abstract, descriptive, noticed, witness, hot, press, study, catalonians, ours, explanation, oppressive, neighbor, defense, peeved, ordered, commander, slaughter, rebellion, capitulation, 11th, privileges, revoked, downhill, digress, surrender, meant, french, loaned, exposing, third, decade, fourth, brief, unsuccessful, attack, bent, eluded, large, exhibit, catastrophe, 1714, largely, 1713, crazy, catalans, kept, candidate, carlos, iii, abandoned, imperial, throne, 1711, becoming, holy, roman, emperor, appears, former, digs, remain, refused, acknowledge, felipe, stayed, death, witnessed, viewing, crypt, vienna, tomb, celebrates, 1706, liberation, kapuzinergruft, iberian, western, front, portuguese, border, animated, holdings, iberia, narrative, forces, managed, repulse, occupations, recapture, territory, battlefield, victory, almansa, obligingly, facilitated, reconquest, choosing, abandon, allies, contingent, captured, brihuega, museo, del, ejército, museu, història, catalunya, relevant, irregular, miquelets, rose, taxes, governance, olivares, home, songs, conflict, anthem, els, segadors, reaper, languedoc, visited, excursion, pyrenean, trouble, pirates, fun, came, origin, legend, senyera, 897, grateful, dipped, fingers, mortally, wounded, blood, dragged, shield, thanking, requested, killers, reward, differently, apparently, dating, 13c, prefer, stories, pen, sword, parchment, quill, ink, indeed, austrians, tale, crimson, streaked, half, waterfront, randomly, selected, destination, covers, region, hear, present, mannequin, cavemen, jump, 19th, rooms, kingdoms, principalities, shake, proverbial, stick, catalonian, speaking, struck, predominance, anglo, bookstores, geoffrey, parker, lynch, elliot, keegan, translations, histories, heavy, seeming, silence, senior, native, imperialism, reality, smartphone, headphone, fated, modernist, casa, bottló, photo, vertigo, anyway, attractions, mandatory, tourist, dollars, visiting, gaudí, standing, narrow, tower, waiting, wanting, contemplate, sagrada, familia, cathedral, tourists, must, proximity, bake, bâtards, wasn, rioting, spanked, liverpool, champions, survived, ugly, american, brooklynite, incident, train, arrived, northern, capital, seemed, sort, disagreement, parts, country, clue, proudly, hung, apartment, balconies, outnumbering, spying, vestiges, conquistadors, visits, castile, second, trip, side, conquerors, conquered, decide, eastern, belligerent, locating, efficiency, oriented, mindset, wherein, collects, ancient, explores, durations, attributes, cool, stuff, accurate, www, gokhan, collate, chart, heart, desires, wants, bureaucracy, inject, entered, factual, approximation, refine, beta, complete, defined, preparing, appeal, skulkers, assist, quixotic, quest, robust, envision, posting, aspects, combats, queries, lights, consist, phrase, describe, alludes, backward, move, sweater, shaved, sweaters, colorful, analogy, reilly, minor, tweaks, collections, newhailes, deane, takes, sucker, oysters, beauty, camps, filename, nest, dictionaries, layers, discussed, basics, 8th, precisely, target, hits, straddling, divide, pair, congratulations, ish, return, pickle, binary, passes, massage, values, efficient, fewer, lines, beginner, dict, comprehension, chained, encoding, splitting, hurdles, neophyte, discovered, pieces, fit, nutshell, within, interact, mileage, vary, 61404, muddled, curly, brackets, odd, equal, loss, care, robustly, relying, kindness, beggars, choosers, happen, reproduction, partial, tells, expanding, bother, rid, extra, stripped, standardized, identified, consists, carriage, distinctive, colons, ultimately, organized, per, appropriate, minutes, aside, examining, swearing, unstructured, consistent, scheme, judicious, unique, proof, concept, kindly, lawrence, smith, extent, mimicking, fidelity, 20th, maastricht, london, confusing, seen, explicitly, reinforces, best, passwords, solution, hone, tagging, transcriptions, requiring, wading, dross, likelihood, distracted, appear, filter, tested, physically, components, organize, old, quite, notecards, documented, journey, modified, applescripts, rudimentary, explode, rescue, tag, archival, begin, isolated, silos, timespan, needles, haystack, unparseable, automates, took, assistance, progress, show, preliminary, sweet, juicy, tantalizingly, reach, impediments, playing, pdfs, meaningful, systematic, slowly, catching, thus, pockets, multi, collaborators, position, sights, lower, taste, fruits, acquired, five, writing, diving, suggest, scan, decent, notepad, wrangler, killer, app, reads, statistics, identifies, rare, unusually, extracts, nationality, age, mentions, tables, graphs, locate, theaters, analyzed, runs, includes, comparisons, assigned, screen, chaining, objects, logic, copacetic, classes, voyant, among, heat, settings, parameters, performed, unable, faithfully, grant, agencies, recipients, submit, replicability, intriguing, rerun, exploration, super, anywhere, turn, extensible, related, automating, ask, drudgery, counting, sorting, revising, whom, hire, gets, burgeoning, classifying, superseded, neural, nets, specified, seeing, syllabus, meeting, removing, holidays, chair, reports, enrollment, assessment, surveys, scheduler, faculty, timeslots, university, departmental, scheduling, requirements, meets, administrative, school, batch, flexibility, mac, arcgis, boring, tasks, fledged, scriptable, portraits, facebook, reconstructing, soundscapes, genealogical, royal, genealogy, 16m, subsets, hathitrust, publications, periods, comparing, cited, affiliations, sciences, publish, ties, citation, networks, bibliometrics, biggie, extraction, cluster, segments, tend, collocated, keywords, sentiment, fuzzy, spelled, prose, phrasings, grammatical, structures, overuse, nltk, spacy, textacy, gensim, embeddings, calendars, tuesday, written, quick, calculation, calendar, datetime, dateutil, convertdate, dateparser, arrow, direction, contents, enter, pyzotero, interactions, sqlite, mysql, sqlite3, believe, triplets, verb, diagram, edges, measure, hubs, spokes, networkx, multiples, projection, whichever, aka, geocode, spatial, teams, gazetteers, georeferenced, geopandas, cartopy, mapping, explored, timeline, load, matplotlib, seaborn, humans, computers, detailed, description, ryan, cordell, hands, late, occurrence, outwork, misocred, mariborough, due, endings, changed, audit, difficult, capitalizations, liked, capitalization, clues, identifying, recognition, ideally, needs, special, handling, critical, bottleneck, idiosyncratic, historically, variant, usage, widely, varying, genre, vocabulary, sometimes, derived, ocring, irregularly, accuracy, rates, 100, ideal, quality, scans, average, scholar, acquire, ecco, paid, cheap, foreign, labor, theirs, social, scientists, born, everything, lowercase, stripping, imperfectly, retain, atomize, tokens, suggestions, welcome, useful, venues, 000s, skill, ted, underwood, jtb, raven, seriously, illustrate
Text of the page (random words):
ares about the art of indexing anymore some are literally headwords with a giant undifferentiated list of several dozen page numbers separated only by commas some don t even provide any kinds of topics only proper nouns of course the index may well die as more works are consumed digitally web scraping automate downloading content from a website either the text on pages entries in a list images of battle paintings from wikipedia or linked files maybe automate the sparql query on historical battles i posted about awhile back or download a bunch of letters from a site that puts each letter on its own separate page maybe automate scraping publication abstracts from a website based off records in zotero with a library like beautifulsoup web form entry i d like to create code that would automate copying bib info from zotero author title date pages etc and then paste it into our library s online ill form in the respective webform fields which of course aren t in the same order as the zotero field order that means a bunch of cutting and pasting for every request look up associated information on an entity person place organization with linked open data e g find the date of birth for person x mentioned in your text document via the web rdflib added by request get it apis numerous institutional websites allow you to access their online data more directly through api rather than using brute force scraping to harvest information from their pages html those websites have apis application program interfaces to allow more sophisticated downloading of information you can search their site for api to see if they offer it requests convert information into data my two previous posts on aha department enrollments and parsing long notes illustrate how this can be done with python code that converts information into data is particularly important low hanging fruit for historians since we lack a lot of already digitized datasets this kind of code allows us to create them with our own data clean dirty ocr text correct ocr errors and generally make a document more readable by humans and computers a good detailed description is ryan cordell s q i jtb the raven taking dirty ocr seriously this requires a lot of hands on work with the code which i ve been doing of late e g find every occurrence of out work and convert it to outwork so we can count them all the same way find every misocred mariborough and convert it to marlborough there are big lists of common errors available to make this search and edit process a bit more precise but since you ll never guess every word that might be hyphenated due to line endings finding every hyphenated word like be siege and converting it back to besiege is easy enough with regular expressions you can even create a list of all the changes your code makes i e create a dictionary with each mistake and what it was changed to if you want to audit the process more difficult are the questions of capitalizations case especially when our 17 18c authors liked to capitalize lots of nouns yet modern nlp uses capitalization as one of its clues for identifying proper nouns named entity recognition ideally you d have this code as a series of functions so that you can run the various corrections across an entire folder of documents based on what that source needs then you could use more code to check for any other errors or documents that require special handling i d argue that this is currently the most critical area for digital history the bottleneck in fact so few historians have their own sources in clean full text yet it s also a very idiosyncratic thing to program based on historically variant word usage and widely varying source genre vocabulary as well as the sometimes random errors derived from ocring irregularly set type from a few hundred years ago it d be great if ocr accuracy rates were 100 but that ideal would seem to require having high quality scans of the originals which your average scholar does not have and will likely never acquire because we don t actually own the originals as a result cleaning historical ocred text is the area with the least amount of pre made code available big projects like eebo and ecco paid cheap foreign labor to type theirs by hand we should also note that lots of social scientists for example talk about preprocessing text but by that they mean standardizing the spelling of text that s been born digital making everything lowercase stripping out punctuation etc historians need a lot of pre preprocessing first because we are dealing with imperfectly ocred text and if we want to retain a cleaned text copy of the original and not just atomize the text string into a list of word tokens then it s even more complicated suggestions welcome ted underwood has provided some useful ideas in various venues dealing with big data 10 000s of texts but his cleaning code on github is a bit above my skill level quantitative analysis load your spreadsheet into the pandas library and run stats plot charts pandas matplotlib seaborn bokeh for interactivity make interactive visualizations create interactive websites with your data and so on bokeh create a visual timeline drawn from data extracted from a document possibly interactive haven t explored this yet but i will mapping quickly make maps including small multiples in whatever projection with whichever features you want to display look up coordinates aka geocode and calculate spatial topological relationships it s good to see that there are several international project teams working on historical gazetteers and a number of groups have georeferenced some early modern maps as well geopandas cartopy network analysis create a network graph diagram of entities and relationships between those entities nodes and edges and measure the topological properties of the network who are the hubs the spokes the nodes networkx relational database interactions with sqlite and mysql sqlite3 i believe you can do the same with graph databases triplets of subject verb object but i haven t really looked at those zotero update and manipulate your zotero records i can already read zotero data into python and hand it off to other libraries for analysis as well as go in the other direction e g mass update fields back in zotero but i d like to create code that will take a pdf page or two of a book s table of contents and enter a separate record for each chapter in the book into zotero automatically adding in all the other book info pyzotero dates automatically calculate duration between two dates convert between os and ns and other historical calendars look up last tuesday s date when mentioned in a letter written on july 7 1700 you could probably automatically insert that date into the source document if desired maybe even do a quick calculation to see how few days you have left on your sabbatical calendar datetime dateutil convertdate dateparser arrow textual analysis this is a biggie for historians create a corpus of texts in your area create a list of people places and events to use for extraction cluster works or segments together by the topic they discuss see how often different authors texts use particular words phrases identify which words tend to be collocated with which other words keywords in context sentiment analysis etc did i mention fuzzy searching finding words that are spelled similarly or finding words that are used in the same context as a given word maybe you want to analyze your own prose which words phrasings grammatical structures do you overuse nltk spacy textacy gensim word embeddings bibliometrics and historiographical analysis from a secondary source extract all the people publications places and time periods dates mentioned and graph map them before comparing them with other authors or analyze the sources cited in the bibliography authors and affiliations years of publication languages etc the sciences have a lot of this already because they mostly publish journals and they re in databases like web of science this also ties into network analysis especially if you want to look at citation networks analyze words phrases from the 16m book hathitrust collection there s a website for that but you can also download the data or subsets at least genealogy parse genealogical data and analyze would be interesting for royal lineages and some work has already been done on that sound analysis haven t played with these but some people are into reconstructing soundscapes and the like image classification and analysis group together all the portraits of person x etc haven t played with these though you have similar classification algorithms in facebook etc lots of full fledged programs are also python scriptable e g both arcgis and qgis have python interfaces which means you can automate many of the boring tasks you need to perform when making more sophisticated maps clean up your computer files batch rename copy delete convert etc with much more flexibility than mac os x s rename function automate lots of administrative school work create a syllabus class schedule that lists the day of week and date for each meeting during the semester removing any holidays or other days off for that specific semester i ll be department chair next year and there are lots of stats and reports on enrollment assessment that i d like to automate collecting the data from databases surveys and then analyze them without me having to manually repeat the entire process every time a computer science colleague will have a student working on a course scheduler next semester given a department s faculty requests the available timeslots and a few dozen university and departmental scheduling requirements come up with a schedule that meets all those criteria or at least the most important ones now we have to do this by hand with excel but still and it s a real pain machine learning ai python is also one of the main languages used for this new burgeoning field for historians that might mean classifying documents and topics but i haven t looked into it enough to think about how it could be used some of the above mentioned libraries might well be superseded by machine learning libraries in the future where things like neural nets figure out their own algorithms without rules being specified by the programmer i think we re already seeing a little bit of that with nlp and those are just a few of the things you can do with python so whatever data related project you can think of there s probably a way to do it in python it s not just automating the things that you find yourself doing on the computer over and over and over and over again just as important what are the research questions that you want to ask especially those that would require a lot of drudgery like counting and sorting and revising thousands of documents any software package that will answer that particular question for you will have its own learning curve and there probably aren t many people whom you could hire to do it for you so you will probably be on your own whatever your question there s likely a way to combine the various python tools together in a way that gets you the desired output but that s not all using code also means you can take whatever output and turn it into the input for another bit of code and so on and so on it is practically infinitely extensible you can rerun your code but change a parameter to see the difference it makes what if exploration is super simple and you can easily change a parameter anywhere in the workflow and continue the rest of your code with the new results when you re all done with your code you can run it on another data set or text or a whole folder full and then you can compare the results when you notice an intriguing pattern in one of your sources you can quickly add another bit of code to explore it then you can look for that pattern in your other documents you will also have a record of your process and method which data you used for which analysis how you cleaned the data which settings and parameters you used the order in which you performed your various steps and so on i m guessing that more than a few historians would be unable to repeat much less explain how exactly they got the results they did how faithfully for example do we record our computer based research workflow some grant agencies are beginning to require recipients submit their data and workflow along with their results replicability could even come to mean something in history this historian s killer app for python is a program that reads in a primary or secondary source from a text file and then the code provides statistics on the words and phrases used identifies rare terms that are unusually common in that document compared to some corpus extracts all the proper nouns mentioned provides a statistical overview of their frequency overall and by section of book looks up information on the people say their nationality age etc then looks up the coordinates of mentioned places and maps them according to some criteria by person who mentions the place by where in the text it is mentioned by what other things are mentioned around that place one output of all this could be tables or graphs of the entities in the text word visualizations and the like another output could be automatically created maps not just maps of any of the above entities but small multiple maps that would locate a variable say siege duration across four different theaters and then another set of small multiple maps that would similarly map the same variable by year instead might as well have it make a heat map while you re at it several groups have already created web versions of some of these features voyant tools among them but with your own code you also end up with all these results in the code itself which can be further analyzed with yet more code then your code runs itself on a bunch of other documents and includes comparisons between documents which texts talk more about place x this works for teaching as well as research imagine if you had a class where you assigned a source had the students analyze it and then put an interactive visualization of the document up on the screen to explore this really wouldn t be that hard i already have almost all of the bits and it s just a question of chaining them all together it will take a while to make sure the objects logic and syntax are all copacetic but hopefully it ll be done in time for classes next fall if you re not sure about diving into python i d suggest you start by getting as many of your sources in digital form as possible scan ocr type then get yourself a decent text editor like notepad or text wrangler bbedit and start learning regular expressions but the more historians we get writing python code the more history specific code we can build off of so let s get started december 3 2018 in methodology 2 comments from historical source to historical data where i offer a taste of just one of ...
|