Meta tags:
description= Genteelly Observing the Enemy since 2011;
Headings (most frequently used words):
the, python, what, exactly, does, for, do, step, in, of, historical, recent, skulking, holes, and, corners, world, siege, sale, cleaning, text, with, sabbatical, rear, view, mirror, you, might, be, millner, from, source, to, data, where, historians, are, 2017, have, mentioned, future, is, digital, early, modern, spain, on, budget, part, reminder, smh, 2019, computer, program, scholars, this, historian, posts, comments, categories, bibliography, archives, blogroll, parsing, unparseable, converting, semi, structured, document, into, files,
Text of the page (most frequently used words):
the (548), and (263), you (251), that (149), with (119), for (114), python (74), but (74), can (70), have (68), from (59), this (58), your (53), all (51), code (50), into (45), text (44), more (44), are (44), data (43), what (39), like (38), some (38), just (35), those (35), will (31), want (31), which (30), one (29), list (28), other (27), use (26), not (26), then (26), there (25), out (25), they (25), create (24), about (24), each (24), history (23), information (23), #digital (22), could (22), document (22), does (21), their (21), few (21), them (21), how (21), over (20), any (20), was (20), well (19), also (19), word (19), already (18), world (18), most (18), has (18), now (17), early (17), modern (17), may (17), spanish (17), been (17), programming (17), make (17), words (17), computer (17), need (16), would (16), lots (16), even (16), own (16), 2018 (15), war (15), siege (15), research (15), when (15), catalan (15), many (15), its (15), historians (15), specific (15), analysis (15), documents (15), maps (14), least (14), bit (14), had (14), entries (14), get (13), 2014 (13), might (13), see (13), through (13), because (13), time (13), lot (13), several (13), libraries (13), historical (12), 2012 (12), august (12), 2013 (12), 2015 (12), 2019 (12), two (12), using (12), another (12), people (12), much (12), things (12), way (12), learning (12), folio (12), online (11), book (11), exactly (11), next (11), our (11), find (11), these (11), long (11), after (11), here (11), should (11), analyze (11), same (11), letters (11), only (11), results (11), different (11), excel (11), december (10), january (10), 2016 (10), together (10), year (10), who (10), barcelona (10), were (10), where (10), every (10), used (10), done (10), his (10), know (10), above (10), probably (10), language (10), clean (10), source (10), don (10), work (10), etc (10), maybe (10), automate (10), web (10), september (9), 2017 (9), zotero (9), day (9), process (9), think (9), back (9), various (9), case (9), separate (9), than (9), since (9), years (9), map (9), important (9), set (9), first (9), between (9), mentioned (9), big (9), available (9), dictionary (9), files (9), item (9), convert (9), format (9), errors (9), july (8), spain (8), search (8), take (8), number (8), off (8), count (8), making (8), authors (8), databases (8), read (8), sure (8), allows (8), example (8), file (8), easily (8), string (8), means (8), date (8), require (8), page (8), website (7), comments (7), skulking (7), holes (7), corners (7), blog (7), march (7), june (7), october (7), sources (7), methodology (7), cleaning (7), new (7), fortress (7), once (7), figure (7), did (7), better (7), access (7), small (7), museum (7), before (7), say (7), course (7), yak (7), full (7), really (7), start (7), future (7), run (7), still (7), structured (7), look (7), journal (7), yet (7), common (7), person (7), numbers (7), comment (6), content (6), wars (6), november (6), february (6), note (6), library (6), great (6), sale (6), recent (6), someone (6), including (6), let (6), sense (6), following (6), very (6), spent (6), particular (6), particularly (6), thing (6), shaving (6), never (6), steps (6), keep (6), check (6), wikipedia (6), whatever (6), order (6), allow (6), change (6), easy (6), value (6), nlp (6), perform (6), changes (6), jupyter (6), notebook (6), notes (6), tools (6), created (6), numerous (6), type (6), project (6), without (6), low (6), ocr (6), program (6), nouns (6), versions (6), features (6), texts (6), based (6), haven (6), enough (6), programs (6), functions (6), started (5), site (5), com (5), name (5), view (5), free (5), bibliography (5), april (5), sieges (5), military (5), google (5), devonthink (5), sabbatical (5), posts (5), via (5), leave (5), around (5), possibly (5), learn (5), doesn (5), ways (5), madrid (5), imagine (5), three (5), end (5), include (5), getting (5), real (5), fact (5), across (5), hand (5), line (5), combine (5), details (5), graph (5), past (5), step (5), yourself (5), campaign (5), little (5), add (5), dates (5), structure (5), pretty (5), entry (5), dozen (5), good (5), parse (5), record (5), copy (5), result (5), automatically (5), bunch (5), place (5), individual (5), thousands (5), historian (5), hanging (5), hard (5), projects (5), places (5), output (5), question (5), machine (5), working (5), letter (5), manipulate (5), types (5), tcp (5), wordpress (4), write (4), 2011 (4), france (4), conference (4), millner (4), finding (4), held (4), part (4), worth (4), visit (4), especially (4), wouldn (4), otherwise (4), fortunately (4), beginning (4), events (4), citadel (4), tell (4), fort (4), 1640 (4), while (4), why (4), give (4), main (4), right (4), said (4), heritage (4), catalonia (4), similar (4), therefore (4), guess (4), instead (4), english (4), city (4), succession (4), whole (4), during (4), four (4), semi (4), names (4), something (4), actually (4), almost (4), come (4), simple (4), scraping (4), open (4), deal (4), basic (4), series (4), doing (4), months (4), key (4), sorts (4), resulting (4), loop (4), delimiter (4), multiple (4), cleaned (4), camp (4), info (4), transcription (4), others (4), looks (4), regular (4), expressions (4), parsing (4), precise (4), problem (4), terms (4), phrases (4), likely (4), put (4), github (4), fruit (4), generally (4), ocred (4), dirty (4), kind (4), form (4), proper (4), entities (4), interactive (4), workflow (4), along (4), mean (4), makes (4), science (4), rename (4), download (4), network (4), made (4), eebo (4), pages (4), index (4), processing (4), books (4), coding (4), scholars (4), 000 (4), email (3), subscribe (3), account (3), rohan (3), british (3), army (3), primary (3), officers (3), taking (3), marlborough (3), art (3), resources (3), professional (3), miscellaneous (3), interested (3), help (3), smh (3), panels (3), coming (3), topic (3), definitely (3), wife (3), less (3), force (3), decided (3), build (3), starting (3), dozens (3), 1705 (3), thought (3), always (3), correct (3), forward (3), difference (3), central (3), recently (3), referendum (3), final (3), performing (3), interesting (3), somebody (3), rest (3), languages (3), ended (3), duke (3), service (3), allied (3), point (3), fall (3), king (3), memory (3), him (3), close (3), aren (3), unlike (3), today (3), currently (3), salses (3), left (3), medieval (3), pain (3), scratch (3), later (3), century (3), again (3), mostly (3), flags (3), far (3), having (3), week (3), last (3), whether (3), sparql (3), battles (3), location (3), post (3), hopefully (3), play (3), try (3), knowledge (3), until (3), qgis (3), reading (3), often (3), eventually (3), reason (3), images (3), journals (3), lists (3), hit (3), bits (3), statistical (3), function (3), returns (3), split (3), methods (3), notice (3), delete (3), standard (3), won (3), editor (3), standardizing (3), regex (3), sample (3), preprocessing (3), nor (3), separated (3), adding (3), top (3), level (3), relational (3), assuming (3), eye (3), sophisticated (3), database (3), searching (3), compared (3), nice (3), days (3), paragraph (3), collection (3), publication (3), compare (3), quickly (3), task (3), textual (3), ones (3), examples (3), image (3), digitized (3), institutional (3), provides (3), corpus (3), variable (3), works (3), syntax (3), repeat (3), folder (3), questions (3), software (3), answer (3), curve (3), rules (3), semester (3), department (3), stats (3), manually (3), entire (3), given (3), both (3), area (3), pdf (3), table (3), websites (3), quantitative (3), edit (3), pre (3), complicated (3), feature (3), natural (3), published (3), humanities (3), behind (3), commands (3), print (3), worry (3), microsoft (3), version (3), packaged (3), vep (3), required (2), log (2), sign (2), subscribed (2), louis (2), smhblog (2), shot (2), lineages (2), archives (2), group (2), society (2), strategy (2), operations (2), graphics (2), fortifications (2), dead (2), conferences (2), battle (2), 18c (2), ememh (2), teaching (2), admin (2), categories (2), jostwald (2), rear (2), mirror (2), older (2), panelists (2), proposals (2), european (2), love (2), considered (2), tramping (2), olympics (2), ditch (2), taken (2), commemorating (2), random (2), overlooking (2), submission (2), reapers (2), revolt (2), centuries (2), human (2), rights (2), torture (2), surprise (2), castle (2), exhibition (2), montjuïc (2), interactivity (2), says (2), opening (2), maintained (2), fortresses (2), below (2), hill (2), park (2), either (2), named (2), fast (2), built (2), decades (2), chapter (2), provide (2), community (2), however (2), government (2), support (2), population (2), against (2), symbol (2), resistance (2), relationship (2), national (2), tourism (2), creating (2), overall (2), saw (2), southern (2), philippe (2), berwick (2), fighting (2), everywhere (2), else (2), besiege (2), rather (2), austrian (2), karl (2), mind (2), clearly (2), earlier (2), familiar (2), minute (2), went (2), tolédo (2), experience (2), display (2), castilian (2), franco (2), thirty (2), keeping (2), track (2), gateway (2), model (2), nationalism (2), away (2), golden (2), video (2), flag (2), prehistory (2), undoubtedly (2), life (2), sticks (2), moderns (2), section (2), john (2), mention (2), guessing (2), speculation (2), historiographical (2), experienced (2), virtual (2), gives (2), moving (2), down (2), impression (2), bread (2), infinitely (2), league (2), looking (2), culture (2), budget (2), europe (2), meantime (2), query (2), listed (2), wikidata (2), curious (2), skills (2), scientist (2), gotta (2), granular (2), https (2), sheets (2), speed (2), times (2), faithful (2), usable (2), dataset (2), spreadsheets (2), warfare (2), downloaded (2), period (2), consider (2), bonus (2), internet (2), goes (2), overview (2), continue (2), yarn (2), surprising (2), allowing (2), oyster (2), extracting (2), depending (2), deep (2), aha (2), keywording (2), cycle (2), parsed (2), hundreds (2), metadata (2), visualization (2), modules (2), become (2), nested (2), items (2), strip (2), strings (2), associated (2), call (2), object (2), single (2), empty (2), seem (2), error (2), programmer (2), focus (2), original (2), easier (2), bbedit (2), formatting (2), seems (2), stage (2), lucky (2), scanned (2), missing (2), changing (2), spelling (2), contractions (2), spaces (2), needed (2), specifically (2), extracted (2), layout (2), towards (2), tabs (2), provided (2), argue (2), subject (2), piece (2), discrete (2), keyword (2), diary (2), drawn (2), being (2), converting (2), explain (2), lack (2), got (2), extract (2), limited (2), focusing (2), offer (2), possible (2), secondary (2), frequency (2), coordinates (2), according (2), criteria (2), visualizations (2), duration (2), similarly (2), groups (2), itself (2), further (2), talk (2), class (2), students (2), explore (2), method (2), pattern (2), parameter (2), input (2), practically (2), package (2), desired (2), field (2), topics (2), looked (2), algorithms (2), schedule (2), collecting (2), colleague (2), student (2), requests (2), interfaces (2), played (2), though (2), classification (2), sound (2), discuss (2), identify (2), context (2), calculate (2), 1700 (2), insert (2), update (2), records (2), mass (2), fields (2), relationships (2), nodes (2), topological (2), properties (2), international (2), visual (2), bokeh (2), spreadsheet (2), pandas (2), plot (2), charts (2), readable (2), requires (2), hyphenated (2), mistake (2), capitalize (2), uses (2), entity (2), corrections (2), hundred (2), ago (2), high (2), originals (2), amount (2), punctuation (2), dealing (2), ideas (2), previous (2), request (2), api (2), brute (2), apis (2), application (2), downloading (2), linked (2), copying (2), paste (2), pay (2), indexed (2), adobe (2), indexing (2), import (2), importance (2), literally (2), giant (2), kinds (2), digitally (2), fantasy (2), custom (2), usually (2), scenes (2), beyond (2), general (2), specialized (2), domain (2), tutorials (2), plenty (2), thanks (2), certain (2), pass (2), block (2), principles (2), attention (2), matter (2), quote (2), perfect (2), algorithm (2), expertise (2), emphasize (2), purchase (2), environments (2), automated (2), control (2), larger (2), upon (2), increase (2), inputs (2), manipulations (2), wish (2), summary (2), nothing (2), lets (2), layer (2), browser (2), variety (2), asking (2), matrix (2), 18th (2), share (2), myself (2), learned (2), geographical (2), summer (2), discoveries (2), installed (2), transcribed (2), england (2), volume (2), countries (2), genteelly (2), observing (2), enemy (2), design, loading, collapse, bar, manage, subscriptions, reader, report, privacy, join, 144, subscribers, quatorze, pike, oderint, dum, probent, perspectives, turenne, anno, domini, 1672, blogroll, wss, weapons, politics, veterans, travel, tactics, religion, recruitment, publishing, prussia, professionalism, pow, ottomans, netherlands, navy, enlightenment, mercenaries, medical, xiv, logistics, laws, italy, italian, ireland, intelligence, gtd, captains, gis, fwr, firearms, finance, engineers, emconventions, crusades, cavalry, caturday, casualties, britain, bcw, austria, artillery, sizes, armor, architecture, 30yw, 9yw, 7yw, uncategorized, historiography, select, category, wayne, lee, pradana, friends, propose, contact, seeking, forum, deadline, panel, paper, putting, alma, mater, ohio, state, columbus, theme, soldiers, civilians, cauldron, although, papers, germane, annual, reminder, caught, hosted, 1992, cue, archer, holdover, team, butts, dry, montjuic, locations, measurements, establish, meter, 1792, plaque, slightly, depressing, facts, occupying, stone, garrison, intimidate, town, coincidence, castilians, broke, bombard, restless, barcelonans, civil, became, violations, imprisonment, execution, political, prisoners, ends, quotation, universal, declaration, article, freedom, seek, involved, 1697, moat, monsters, guarding, ravelin, nope, guérite, smell, urine, photos, talking, zoom, descriptions, underwent, street, splay, embrasures, outward, funnel, themselves, flower, beds, aerial, bearings, complex, screenshot, surrounded, cable, car, funicular, walk, interest, sites, pronounced, mont, jew, initial, jewish, burial, ground, outcropping, mediterranean, transformed, lighthouse, successive, formidable, ethnographic, autonomous, northeast, known, implicitly, recognised, 1978, constitution, disagreements, tax, situation, hostility, independence, soared, polarising, outcome, remains, uncertain, previously, ruling, party, rebranded, non, binding, symbolic, pressure, politically, sensitive, significance, strained, johannes, venetia, identity, observations, edited, catherine, palmer, jacqueline, tivers, routledge, abstract, descriptive, noticed, witness, hot, press, study, catalonians, ours, explanation, oppressive, neighbor, defense, peeved, ordered, commander, slaughter, rebellion, capitulation, 11th, privileges, revoked, downhill, digress, surrender, meant, french, loaned, exposing, third, decade, fourth, brief, unsuccessful, attack, bent, eluded, large, exhibit, catastrophe, 1714, largely, 1713, crazy, catalans, kept, candidate, carlos, iii, abandoned, imperial, throne, 1711, becoming, holy, roman, emperor, appears, former, digs, remain, refused, acknowledge, felipe, stayed, death, witnessed, viewing, crypt, vienna, tomb, celebrates, 1706, liberation, kapuzinergruft, iberian, western, front, portuguese, border, animated, holdings, iberia, narrative, forces, managed, repulse, occupations, recapture, territory, battlefield, victory, almansa, obligingly, facilitated, reconquest, choosing, abandon, allies, contingent, captured, brihuega, museo, del, ejército, museu, història, catalunya, relevant, irregular, miquelets, rose, taxes, governance, olivares, home, songs, conflict, anthem, els, segadors, reaper, languedoc, visited, excursion, pyrenean, trouble, pirates, fun, came, origin, legend, senyera, 897, grateful, dipped, fingers, mortally, wounded, blood, dragged, shield, thanking, requested, killers, reward, differently, apparently, dating, 13c, prefer, stories, pen, sword, parchment, quill, ink, indeed, austrians, tale, crimson, streaked, half, waterfront, randomly, selected, destination, covers, region, hear, present, mannequin, cavemen, jump, 19th, rooms, kingdoms, principalities, shake, proverbial, stick, catalonian, speaking, struck, predominance, anglo, bookstores, geoffrey, parker, lynch, elliot, keegan, translations, histories, heavy, seeming, silence, senior, native, imperialism, reality, smartphone, headphone, fated, modernist, casa, bottló, photo, vertigo, anyway, attractions, mandatory, tourist, dollars, visiting, gaudí, standing, narrow, tower, waiting, wanting, contemplate, sagrada, familia, cathedral, tourists, must, proximity, bake, bâtards, wasn, rioting, spanked, liverpool, champions, survived, ugly, american, brooklynite, incident, train, arrived, northern, capital, seemed, sort, disagreement, parts, country, clue, proudly, hung, apartment, balconies, outnumbering, spying, vestiges, conquistadors, visits, castile, second, trip, side, conquerors, conquered, decide, eastern, belligerent, locating, efficiency, oriented, mindset, wherein, collects, ancient, explores, durations, attributes, cool, stuff, accurate, www, gokhan, collate, chart, heart, desires, wants, bureaucracy, inject, entered, factual, approximation, refine, beta, complete, defined, preparing, appeal, skulkers, assist, quixotic, quest, robust, envision, posting, aspects, combats, queries, lights, consist, phrase, describe, alludes, backward, move, sweater, shaved, sweaters, colorful, analogy, reilly, minor, tweaks, collections, newhailes, deane, takes, sucker, oysters, beauty, camps, filename, nest, dictionaries, layers, discussed, basics, 8th, precisely, target, hits, straddling, divide, pair, congratulations, ish, return, pickle, binary, passes, massage, values, efficient, fewer, lines, beginner, dict, comprehension, chained, encoding, splitting, hurdles, neophyte, discovered, pieces, fit, nutshell, within, interact, mileage, vary, 61404, muddled, curly, brackets, odd, equal, loss, care, robustly, relying, kindness, beggars, choosers, happen, reproduction, partial, tells, expanding, bother, rid, extra, stripped, standardized, identified, consists, carriage, distinctive, colons, ultimately, organized, per, appropriate, minutes, aside, examining, swearing, unstructured, consistent, scheme, judicious, unique, proof, concept, kindly, lawrence, smith, extent, mimicking, fidelity, 20th, maastricht, london, confusing, seen, explicitly, reinforces, best, passwords, solution, hone, tagging, transcriptions, requiring, wading, dross, likelihood, distracted, appear, filter, tested, physically, components, organize, old, quite, notecards, documented, journey, modified, applescripts, rudimentary, explode, rescue, tag, archival, begin, isolated, silos, timespan, needles, haystack, unparseable, automates, took, assistance, progress, show, preliminary, sweet, juicy, tantalizingly, reach, impediments, playing, pdfs, meaningful, systematic, slowly, catching, thus, pockets, multi, collaborators, position, sights, lower, taste, fruits, acquired, five, writing, diving, suggest, scan, decent, notepad, wrangler, killer, app, reads, statistics, identifies, rare, unusually, extracts, nationality, age, mentions, tables, graphs, locate, theaters, analyzed, runs, includes, comparisons, assigned, screen, chaining, objects, logic, copacetic, classes, voyant, among, heat, settings, parameters, performed, unable, faithfully, grant, agencies, recipients, submit, replicability, intriguing, rerun, exploration, super, anywhere, turn, extensible, related, automating, ask, drudgery, counting, sorting, revising, whom, hire, gets, burgeoning, classifying, superseded, neural, nets, specified, seeing, syllabus, meeting, removing, holidays, chair, reports, enrollment, assessment, surveys, scheduler, faculty, timeslots, university, departmental, scheduling, requirements, meets, administrative, school, batch, flexibility, mac, arcgis, boring, tasks, fledged, scriptable, portraits, facebook, reconstructing, soundscapes, genealogical, royal, genealogy, 16m, subsets, hathitrust, publications, periods, comparing, cited, affiliations, sciences, publish, ties, citation, networks, bibliometrics, biggie, extraction, cluster, segments, tend, collocated, keywords, sentiment, fuzzy, spelled, prose, phrasings, grammatical, structures, overuse, nltk, spacy, textacy, gensim, embeddings, calendars, tuesday, written, quick, calculation, calendar, datetime, dateutil, convertdate, dateparser, arrow, direction, contents, enter, pyzotero, interactions, sqlite, mysql, sqlite3, believe, triplets, verb, diagram, edges, measure, hubs, spokes, networkx, multiples, projection, whichever, aka, geocode, spatial, teams, gazetteers, georeferenced, geopandas, cartopy, mapping, explored, timeline, load, matplotlib, seaborn, humans, computers, detailed, description, ryan, cordell, hands, late, occurrence, outwork, misocred, mariborough, due, endings, changed, audit, difficult, capitalizations, liked, capitalization, clues, identifying, recognition, ideally, needs, special, handling, critical, bottleneck, idiosyncratic, historically, variant, usage, widely, varying, genre, vocabulary, sometimes, derived, ocring, irregularly, accuracy, rates, 100, ideal, quality, scans, average, scholar, acquire, ecco, paid, cheap, foreign, labor, theirs, social, scientists, born, everything, lowercase, stripping, imperfectly, retain, atomize, tokens, suggestions, welcome, useful, venues, 000s, skill, ted, underwood, jtb, raven, seriously, illustrate
Text of the page (random words):
ch questions that you want to ask especially those that would require a lot of drudgery like counting and sorting and revising thousands of documents any software package that will answer that particular question for you will have its own learning curve and there probably aren t many people whom you could hire to do it for you so you will probably be on your own whatever your question there s likely a way to combine the various python tools together in a way that gets you the desired output but that s not all using code also means you can take whatever output and turn it into the input for another bit of code and so on and so on it is practically infinitely extensible you can rerun your code but change a parameter to see the difference it makes what if exploration is super simple and you can easily change a parameter anywhere in the workflow and continue the rest of your code with the new results when you re all done with your code you can run it on another data set or text or a whole folder full and then you can compare the results when you notice an intriguing pattern in one of your sources you can quickly add another bit of code to explore it then you can look for that pattern in your other documents you will also have a record of your process and method which data you used for which analysis how you cleaned the data which settings and parameters you used the order in which you performed your various steps and so on i m guessing that more than a few historians would be unable to repeat much less explain how exactly they got the results they did how faithfully for example do we record our computer based research workflow some grant agencies are beginning to require recipients submit their data and workflow along with their results replicability could even come to mean something in history this historian s killer app for python is a program that reads in a primary or secondary source from a text file and then the code provides statistics on the words and phrases used identifies rare terms that are unusually common in that document compared to some corpus extracts all the proper nouns mentioned provides a statistical overview of their frequency overall and by section of book looks up information on the people say their nationality age etc then looks up the coordinates of mentioned places and maps them according to some criteria by person who mentions the place by where in the text it is mentioned by what other things are mentioned around that place one output of all this could be tables or graphs of the entities in the text word visualizations and the like another output could be automatically created maps not just maps of any of the above entities but small multiple maps that would locate a variable say siege duration across four different theaters and then another set of small multiple maps that would similarly map the same variable by year instead might as well have it make a heat map while you re at it several groups have already created web versions of some of these features voyant tools among them but with your own code you also end up with all these results in the code itself which can be further analyzed with yet more code then your code runs itself on a bunch of other documents and includes comparisons between documents which texts talk more about place x this works for teaching as well as research imagine if you had a class where you assigned a source had the students analyze it and then put an interactive visualization of the document up on the screen to explore this really wouldn t be that hard i already have almost all of the bits and it s just a question of chaining them all together it will take a while to make sure the objects logic and syntax are all copacetic but hopefully it ll be done in time for classes next fall if you re not sure about diving into python i d suggest you start by getting as many of your sources in digital form as possible scan ocr type then get yourself a decent text editor like notepad or text wrangler bbedit and start learning regular expressions but the more historians we get writing python code the more history specific code we can build off of so let s get started december 3 2018 in methodology 2 comments from historical source to historical data where i offer a taste of just one of the low hanging fruits acquired over my past five months of python the sabbatical digital history is slowly catching on but thus far my impression is that it s still limited to those with deep pockets big multi year research projects with a web gateway and lots of institutional support including access to computer scientist collaborators since i m not in that kind of position i ve set my sights a bit lower focusing on the low hanging fruit that s available to historians just starting out with python yet much of this sweet juicy low hanging fruit is tantalizingly still just out of reach undoubtedly you already know that one of the big impediments to digital history generally and to historians playing with the python programming language specifically is the lack of historical sources in a structured digital format we ve got thousands of image pdfs even ocred ones but it s hard to extract meaningful information from them in any structured way and if you want to clean that dirty ocr or analyze the text in any kind of systematic way you need it digitized but in a structured format my most recent python project has been to create some python code that automates a task i m sure many historians could use parsing a big long document of textual notes documents into a bunch of small ones it took one work day to create it without the assistance of my programming wife so i know i m making progress eventually i ll clean the code up and put it on my github account for all to use but for now i ll just explain the process and show the preliminary results for examples of how others have done this with python check out the programming historian particularly this one parsing the unparseable converting a semi structured document into files if you re like me you have lots of historical documents most numerous are the thousands of letters diary and journal entries from dozens of different authors each collection of documents is likely drawn from a specific publication or archival collection which means they begin being all isolated in their little silos if you re lucky they re already in some type of text format ms word or excel a text file what have you and that s great if you want just to search for text strings or maybe even use regular expressions but if you want more if say you want to compare person a s letters with person b s letters over the same timespan or compare what they said about topic x or what they said on date z then you need to figure out a way to make them more easily compared to quickly and easily find those few needles in the haystack the time tested strategy for historians has been to physically split up all your documents into discrete components and keyword and organize those individual letters or diary entries or in the old days which are still quite new for some historians you d use notecards i ve already documented my own research journey away from word documents to digital tools see devonthink tag i even created modified a few applescripts to automate this very problem in devonthink in a rudimentary way one for example can explode i e parse a document by creating a new document for every paragraph in the starting document nice but it can be better python to the rescue the problem lots of text files of notes and transcriptions of letters but not very granular and therefore not easily compared requiring lots of wading through dross with the likelihood of getting distracted this is particularly a problem is you re searching for common terms or phrases that appear in lots of different letters wouldn t it be nice if you could filter your search by date or some other piece of metadata the solution use python code to parse the documents say individual letters or entries for a specific day into separate files making it easy to hone in on the precise subject or period you re searching for as well as precise tagging and keywording step 1 for proof of concept i started with a transcription of a campaign journal kindly provided me by lawrence smith in a word document i m sure you have dozens of similar files he was faithful in his transcription even to the extent of mimicking the layout of the information on the page with the use of tabs spaces and returns great for format fidelity but not great for easily extracting important information particularly if you want for example june to be right next to 20th instead of on the line below separated by a bunch of officers names maastricht and london are actually a bit confusing because i m pretty sure the place names after the dates are that day s passwords at least that s what i ve seen in other campaign journals that some of the entries explicitly list a camp location reinforces my speculation of course people can argue about which information is important which is yet another reason why it s best if you can do this yourself aside as you are examining the layout of the document to be parsed you should also have one eye towards the future in this case that means swearing to yourself that i will never again take unstructured notes that will require lots of regex for parsing in other words if you want to make your own notes usable by the computer and don t already have a sophisticated database set up for data entry use a consistent format scheme across sources that is easy to parse automatically for example judicious use of tabs and unique formatting step 2 clean up the text specifically make the structure more standardized so different bits of info can be easily identified and extracted for this document that means making sure each first line only consists of the date and camp location when available that each entry is separated by two carriage returns and adding a distinctive delimiter in this case two colons between each folio because you ll ultimately have the top level of your structured data organized by folio with entries multiple entries per folio this is a one to many relationship for those of you familiar with relational databases like access cleaning the text can be easily done with regex allowing you to cycle through and make the appropriate changes in minutes assuming you know your regular expressions that is the result looks like this note that this stage is not changing the content i e it s not preprocessing the text doing things like standardizing spelling or expanding contractions or what have you nor did i bother getting rid of extra spaces etc those can be stripped with python as needed for this specific document note as well that some of the formatting for the officers of the day is muddled the use of curly brackets seems odd which might equal loss of information but if that info s important you should take care to figure out how to robustly record it at the transcription stage if you re relying on the kindness of others beggars can t be choosers but if you re lucky you happen to have a scanned reproduction of a partial copy of this journal from another source which tells you what information might be missing from the transcription camp journal sample of above from british library add ms 61404 f 45 you probably could do this standardizing within your python code in jupyter notebook but i find it easier to interact with regex in my text editor bbedit your mileage may vary step 3 once you get the text in a standard format like the above you read it into python and convert it into a structured data set if you don t know python at all the following details won t make sense so go read up on some python one of the big hurdles for the neophyte programmer as i ve discovered over and over is to see how the different pieces fit together into a whole so that s what i ll focus on here in a nutshell the code does the following after you ve cleaned up the structure of the original document in your text editor read the file into memory as one big long string perform any other cleaning of the content you want then you perform several passes to massage the string into a dictionary with a nested list for the values there may be a better more efficient way to do this in fewer lines but my beginner code does it in three main steps convert the document to a list splitting each item at the f delimiter now you have a list with each folio as a separate item always look at your results for some reason the first item of the resulting list is empty it doesn t seem to be an encoding error so just delete that item from the list before moving on now read the resulting list items into a python dictionary with the dictionary key the folio number and all of the entries on the folio as the value of that folio use the as the delimiter here with the following line of code a comprehension as they call it notice how the strip and split methods are chained together performing multiple changes on the item object in that single bit of code now you use a for loop to parse each value into separate list items using the other delimiter of n n two returns between entries using the string of the value since otherwise it s a list item and the strip and split methods only work on strings this gives you a dictionary with the folio as the dict key and the value is now a nested list with each of the entries associated with its folio as a separate item as you can see with folio 40 s four entries that s pretty much it now you have a structure for your text congratulations your text has become data or data ish at least the resulting python dictionary allows you to search any folio and it will return a list of all the letters entries on that folio you can loop through all those entries and perform some function on with them so that s a good thing to pickle i e write it to a binary file so that it can be easily read back as a python dictionary later on once you have your data structured and maybe add some more metadata to it you can do all sorts of analysis with all of python s statistical nlp and visualization modules but if you are still straddling the devonthink python divide like i am then you ll also want to make these parsed bits available in devonthink add a bit of code to write out each dictionary key value pair to a separate file and you end up with several hundreds of files each file will have only the content for that specific entry making it easy to precisely target your search and keywording the last thing you want to do is cycle through several dozen hits in a long document for that one hit you re actually looking for that s it entry of may 8th 1705 in its own file the beauty is that you can add more to the code try extracting the dates and camps change what information you want to include in the filename etc depending on the structure of the data you re using you might need ...
|