The Historic Strategy Games That Built My Book

For thirty years, strategy game players have been reckoning with the harsh reality that a computer might be able to play a game better than them. Beginning in 1997 with Kasparov vs Deep Blue and ending with Lee Se-Dol vs AlphaGo, AI inched ahead of human performance year by year, culminating in their total victory.

I love that tension, the open question that floats in the air with every game, ‘Can humanity win?’. Every victory and every defeat carried enormous weight. It’s the heart of my novel, The Human Countermove, strategy games and the fight against a mentally superior enemy.

The challenge with writing a strategy book is creating strategies that feel authentic and clever. The kind of ideas that are convincingly grandmaster in skill, but understandable to the general public. In order to achieve that, I had to learn from the best.

Kasparov vs Deep Blue (1997)

This game is the seed at the center of my book. The tipping point for humanity, the moment we realized computers could out-think people. In 1996, Kasparov won 4-2.

In 1997, they had a rematch, Deep Blue won 3.5-2.5.

Those two matches record the exact year engineering overtook training.

My favorite moment from the 1997 match comes in game 2, when Kasparov accused the Deep Blue team of cheating by having a Grandmaster help with a move. Even a computer can get illegal assistance from time-to-time it seems.

But the conflict of the moment is what really captures me. On the one hand, we want to believe a person is capable of outperforming a computer. On the other, what an incredible feat it is to reproduce the mind of a genius with a bit of code and training. Caught in between, the audience cheers both sides, athletic feat against human ingenuity.

Kasparov has a list of mistakes he says he regrets about that match. Moments he could have snatched a draw from a defeat, a victory from a stalemate. The thing is, if he had won, all it would have done is stall the inevitable. Instead of discussing the 1997 Kasparov vs Deep Blue match, we’d be discussing the 1998 Kasparov vs Deep Blue match.

It’s all of this I try to capture in my book. The tension, the conflict, the regret, and the determination to beat the unbeatable.

Now when Chess Engines and AI models face off against one another, they are a tier beyond our best players. A mentor for grandmasters like Magnus Carlsen, and something beyond the rest of our comprehension.

The Opera Game (1858)

This is a lighter game. The Opera game was played by Paul Morphy and The Duke of Brunswick over a century ago. It’s one I draw inspiration from in my novel not as a strategic tool, but as a piece of chess culture. The Opera Game represents the beginning of a chess student’s education, one of the very first games a novice will be introduced to.

Paul Morphy makes strong, understandable decisions against a much weaker opponent, rapidly gains the advantage, and wins in style. But it’s not just a game, it’s a story. The best in the world dragged into the Duke’s box to play a chess game in the middle of an opera. For beginners, it weaves a romance around chess, and attaches a narrative to one of their first lessons.

In my book, the protagonist Zouk does a lot of teaching on the side, as many professional players find themselves doing. When an opportunity to lecture to a big audience comes around and he realizes the inexperience of his listeners, he abandons the esoteric analysis had prepared, and leans on a tried and true classic with a fun story, The Highway Game.

Go: Lee Se-Dol vs AlphaGo (2017)

Lee Se-Dol vs AlphaGo ended in a 1-4 result. For those of us that had been tracking the development of computers since Deep Blue’s game against Kasparov, seeing AlphaGo take its victory wasn’t a surprise. Go is much more computationally difficult than chess, but Moore’s Law is a powerful force.

But did you notice the scoreboard? Lee Se-Dol won the fourth game. That was an upset.

Against Google’s best engineers and decades of neural networking and algorithmic design, a human being managed to snatch victory, and it all came from a single move. Move 78.

That move has been gone over, analyzed, and studied for years. It’s believed Move 78 pushed the game into a uniquely complicated position, a position AlphaGo couldn’t calculate. A blind spot in the computer’s play that drew out blunder after blunder.

Lee Se-Dol was like a grandmaster Quality Assurance tester, noticing where AlphaGo was weak and pushing it further and further down that path until its behavior was sub-par. Basically, Lee Se-Dol found a bug.

Even when it seemed impossible, a person beat the unbeatable.

The Hippo and Various Anti-AI Strategies

Since Kasparov vs Deep Blue, a thousand Chess engines have burst onto the scene. Anyone willing to run a bit of code on their computer and risk getting banned can play like a grandmaster. To beat such unsavory characters, grandmasters have had to develop a special set of tools. First and foremost is time.

Consider two games. One gives each player an hour to make all their turns, the other gives each player a minute to make their turns. The first game is deeply thought out, with strong moves that remove all chances of counterplay. The second is superficial, moves borne more from training than thought.

In tight time controls, using a chess engine becomes a liability. The grandmaster can play from their subconscious, but the cheater is stuck waiting for the ‘perfect answer’ from the machine.

Thus we meet The Hippo. The Hippo slows the game down to a crawl. Pieces only move forward a square or two, then build a near-impenetrable fortress. As the opponent approaches, the grandmaster makes every effort to close down the position, keeping the number of moving pieces to a minimum.

With each move, the cheater loses a little more time, and the walls close tighter around them.

As their time dwindles, the cheater is forced to throw in a few of their own moves. These usually turn out to be of a significantly lower quality than what a chess engine can put out. Once the grandmaster has stripped the cheater of their chess engine, they unravel all the complexity of The Hippo and go in for the kill.

Once again, complexity and time as weapons to beat an overthinking machine.

The Battle of Cannae and Real-Time Strategy Games

I love real-time strategy games. The feeling of making a plan, facing the hard truths of reality, making adjustments, and turning the battle in your favor is exhilarating. And they’re so different from a game like Chess or Go. In Chess and Go, the entire shape of the board is transformed in a single move. 

In Real-Time Strategy games like Starcraft, you’re making a new move every second, and it’s only when you add all those little decisions up that you end up with a result.

And in games like that, there’s one particular battle result that everyone is chasing.

During the Second Punic Wars, Hannibal faced a much larger Roman force and turned the battle completely in his favor. The trick? Draw the enemy in, encircle them completely, then tighten the trap.

The game in my book, LINE, isn’t like Chess or Go. It’s a little more practical in nature. In theory, the game is playable on a field, not that most people would enjoy the feeling of being shot by a rubber bag. Because of the practical realities of squadrons facing off against one another, tactics like Chess’ fork and pin don’t translate.

But what does translate, is the greatest military trap of all time. Let the enemy over-extend themselves, wait for the right moment, and strike.

Final Words

There are plenty of other strategy games I no doubt pulled inspiration from. Things like the Total War games, Role Playing Games, X-COM, but Chess was my guiding star. It’s funny, once you open your mind to a question like, ‘how does a person beat an AI in strategy?’, you realize how many other people already pondered the same question. 

AIs have been kicking Mankind’s collective butt for thirty years. It’s nearly impossible to imagine a person turning it around on them. But nearly impossible is still possible, it only takes the right person and the right techniques to turn things around. Even when the robot brains out-think us on every front, we can still squeak out a victory every now and then. Especially when we’re learning from everything that’s available.

In The Human Countermove, my protagonist Zouk Solinsen is the right person with the right techniques. The skills to outsmart computational genius.

My debut novel, THE HUMAN COUNTERMOVE is now available for purchase!

My Debut Novel “The Human Countermove” Is Now Available!

I’m gonna keep this update brief.

After three years my debut novel is now available to purchase on Amazon! It’s a cerebral, near-future sci-fi built from my love of strategy games like Chess. In the next few days I will be releasing a post discussing all the different strategies and games I built my book from, but today it’s all about the celebration!

Thank you to all my readers, my family, and my friends. Becoming a novelist took a lot longer than I expected, but I’ve enjoyed every little project along the way. The terrible practice novel, the staged reading of my play, the years developing ed-tech stories for students, each project was a step on my journey here.

Don’t worry, I have no intention of stopping. My next project (Project APHELION) is already about 10% of the way through its second draft, so hopefully it won’t be too long before we’re back here again with another exciting story.

As I schedule appearances at book signings, farmer’s markets, and reader events, I will post them here.

Thank you again, and happy reading.

– Logan Sidwell

The Human Countermove is now available for purchase! https://www.amazon.com/dp/B0FM9R7T5F

In a nation ruled by AI Minds, productivity is everything—even play.

Once a legend in the world of strategy games, Zouk Solinsen is now just another burnout in a society obsessed with efficiency. But when the Minds announce a high-stakes tournament—with a seat on the ruling council as the prize—Zouk is drawn back into the fray, determined to reshape the future.

With help from the enigmatic Torrez Institute, Zouk racks up early victories against the Minds. But when Maya Torrez reveals the cost of her support—a violent coup against the Minds—he rejects it and strikes out alone.

Now, with no allies, dwindling resources, and a nation on the brink, Zouk faces the biggest game of his life—and a final, impossible choice: reform the system from within, or burn it all down.

Three Years Later… I Have a Novel

On September 1st, my debut novel is being released. The Human Countermove. Getting it released is incredibly exciting, and knowing it took three years fills me with a quiet dread. The journey has been incredibly long. Two years to write it. One year to decide what to do with it, and now it’s available for sale. I’m counting every pre-order on a little calendar, crossing off a square with every sale.

Not that you can trust me, but it’s my opinion I’ve written a compelling book. My mom liked it for one. That’s a big improvement over my practice novel. My beta readers liked it, I even managed to convince one of my readers to review two different drafts, which is unheard of in the beta reader space. Usually you only get one chance to impress someone.

But it’s here. It’s been professionally edited, copy-edited, and gone over again and again. Ready for scrutinizing eyes.

The Journey

They say the first one million words are practice. I believe I hit the equivalent of one million words somewhere near the end of my first draft. There was a day when a switch flipped in my head. From then on, my understanding of scene composition, dialog, and character motivations was just, clearer.

For someone editing their first book, a sudden jump in skill is very bad news. It meant I had to face my rough, rough, rough first draft and clean it up with a newfound understanding of storytelling. That’s a lot of work for a single broom.

I lost momentum a couple of times. My systems for reliably writing weren’t in place yet. One weekend I’d pump out 13,000 words, then nothing for a month.

Even the soul of the story wasn’t there on the first go-around. I found it partway into the second draft. A great idea that really clarified the narrative. Funny enough, I wanted to put that heart in the sequel. My editor talked me out of it, convinced me that good ideas are meant to be spent, and that my debut should be as strong as it can be.

In my opinion the back third of this book is where it excels, a final arc that imbues the whole story with purpose. The place where all those funny little ideas were vacuumed out of a hypothetical sequel and pulled into the original.

Choosing to Self Publish

I’m an impatient man. It’s silly of me to be impatient after spending two years writing up a draft, but I was ready for this project to be out there. I’ve met plenty of writers sitting on twelve novels just waiting for the right agent to turn them into stars, that’s not the path for me.

The scariest part of self-publishing is knowing that every inch of success is entirely on you. That also means if the book only sells a dozen copies, it’s your fault. For me, that didn’t seem so bad. I’d rather improve by releasing my work and letting people give me honest feedback than hide away and write book after book on my own. I’ve never worked on something that didn’t get released to the public within four months of being finished before, so a year of waiting was an eternity.

Now that the time is here, I’m really enjoying the process. Soon there will be something out in the world that I’m proud of, something I made, something I’m eager to share. Lately I’ve been attending a lot of farmer’s markets. I haven’t made a single sale, but the experience has been a blast. I get to spend time speaking to real people, giving advice to novice writers, learning what different readers like reading. After all this time on my own pushing to finish a product, getting to know someone else’s story is sort of, healing.

My review of self-publishing so far: Owning my own book and owning my own success is hard work and an absolute joy.

The Novel

I can’t write this whole thing up without talking about my novel! The book is titled “The Human Countermove”, there’ll be a link and description down at the end. But here, in this little blog, I want to give a more informal description.

The book tells the story of Zouk, a washed up strategy game grandmaster who challenges the three AI rulers of his society to determine society’s future.

It’s a cerebral near-future sci-fi, inspired by my love of chess and strategy games. The premise is drawn from the famous chess match Kasparov vs Deep Blue (1997), where mankind’s best chess player was soundly defeated by an algorithm.

I wrote this thing on the hunt for some narrative payback. In real life, we got our butt handed to us. In The Human Countermove, the big question at the start of the book is, ‘What can a person do to out-think something that is cognitively superior’? Zouk Solinsen is my very own John Henry the steel-driving man, except this time instead of trying to beat the machine by brute-force, Zouk pulls every trick in the book to get an advantage.

One thing I fought hard to keep in the book was a rejection of the normal dystopian tropes. So often in these things society is irredeemable, and it all descends into war and destruction. The reader watches the conflict between robots and humans pave a fiery trail for centuries, they see the last few untracked humans turn into a rebellion. I’m ready for something new.

Our main character is a victim of a broken system. A system that demands efficiency in every act. Work and play and rest, all measured, all prescribed in particular doses. It’s not unreasonable to be angry. A broken system needs change. But at the heart of the story is one issue, does the system need to be burned down, or do we not yet understand it? Is there something inherently wrong with a society run by AI Minds? Maybe. Or maybe there’s just a separation between what mankind asks for and what we really want.

Conclusion

As silly as it is, I’ve often defined whether or not I’m a writer by the absence of a published book. I’ve worked professionally in the field, I’ve written for graphics teams, voice actors, education companies, by all means, I am a writer. But this was the last hurdle. As soon as this book comes out, I get to say it to myself and mean every word.

Next week, I will be a novelist.

My debut novel is now available for pre-order. Release Date September 1st. I’m still working out the last few kinks on the paperback side, but that option should be made available soon.

https://www.amazon.com/dp/B0FM9R7T5F

In a nation ruled by AI Minds, productivity is everything—even play.

Once a legend in the world of strategy games, Zouk Solinsen is now just another burnout in a society obsessed with efficiency. But when the Minds announce a high-stakes tournament—with a seat on the ruling council as the prize—Zouk is drawn back into the fray, determined to reshape the future.

With help from the enigmatic Torrez Institute, Zouk racks up early victories against the Minds. But when Maya Torrez reveals the cost of her support—a violent coup against the Minds—he rejects it and strikes out alone.

Now, with no allies, dwindling resources, and a nation on the brink, Zouk faces the biggest game of his life—and a final, impossible choice: reform the system from within, or burn it all down.

My First Draft Took 7 months, Here’s What I Learned

I just finished the first draft of my second book. It took 7 months. The final word count was about 87,000 words. That averages out to about 410 words per day. But that’s not the reality.

The reality is half my book was written across 7 very productive weeks, and half my book was written across 5 very unproductive months. Here’s what I learned.

Find The Process

Last week I wrote a post about my writing process. On days I wrote, I always hit my wordcount goal of 1,200 words. But for a long time, getting my butt in the chair and focussing enough to work proved impossible. Then I started pre-writing with a pen and paper, and I put a time on my phone each day for writing and everything got easier.

From the moment I found my process, my average word-count per week shot up to 6,000. About 5.5 days per week on and off. If I had hit that number from the start, the book would have been done in 2 and a half months.

Momentum is Everything

Forming a consistent rhythm is hard. And sometimes life forces us to make exceptions. Here’s what I’ve learned about myself.

If I take a one day break from writing, I can get back to writing the next day without any issue.

If I take a two day break, I get kinda anxious and starting again becomes a challenge.

After three days, the momentum is gone, and I have to start cold.

The next time I’m writing the first draft of a book, I plan on allocating three dedicated months, with only brief weekend retreats to break things up. Once the habit is formed, it’s harder to break it than to fulfill it. But if I give myself too many excuses, too many easy outs, the habit dies before it’s formed.

Love (With Your Novel) is Fleeting

It’s easy to fall in love with a book. It’s much harder to stay in love. You can only work on the same task for so long before you start to hate it. A terrible kind of insecurity bubbles up, a voice in your ear whispers that your story is terrible.

About 3 months into my drafting, I stopped loving my book. Worse, I stopped liking it. And once that happened, getting words on the page was almost impossible.

The good news is: It’s fixable. It took a little wine and dining, but with the right attitude and a careful approach, I was able to rediscover my passion at least twice while getting the thing written.

The process was pretty simple, when I had been away from my book for a couple weeks and the spark was gone, I’d revisit the book the way I had at the start. Begin by visualizing the world, the aesthetics, the look and wonder of the story. The joy of the concept rather than the pain of the details. Then I’d see my characters, the protagonist with all their flaws, and everything they were trying to do. But it was more than seeing them, it was seeing what was still in store for them. I’d have a third of a book written, and I’d be able to look into the future and know what was still on its way. The end of the arc, still not on the page. My love would reignite, I had seen everything I loved about the story and everything that was still in store. It’s the reason I’m telling the story, the idea that bubbles in my stomach and warms my heart.

Too Much Buildup is Bad for The Writer

Ideas are made to be spent. Once they come into your brain, they fill a space of it until the day you get them onto the page. Worse, a great idea likes to return again and again, occupying most of your thoughts as you imagine the same scene from a hundred different angles.

The trouble with all that thinking is the buildup. At the end of the day, you only get to tell the story one way. And what does that mean for all those other perspectives? They’re tossed in the bin. Maybe I get to pull an idea or two along the way, but most of it is just wasted brainspace.

My brain knows it’s wasted work, and it hates it.

If I love a scene too much, my brain does everything in its power to keep me from writing it. To write is to commit, it takes the infinite possibility and beauty of a concept and turns it into concrete words.

For me, the best thing I can do with a scene I love is get through it as soon as possible. Keep the reimaginings low, keep the ways to spruce things up limited, and let the scene be like you saw it for the first time in your head, even when sometimes it’s just two characters chatting in a garage. It’s much easier to edit a poorly written chapter than fill a blank page.

The Outline is Key

My outline was my most important ingredient, it turned the impossible journey of 100,000 words into a bunch of 1,200 word slices. When I lost momentum, I put a list on the wall, a series of individual scenes pulled from my outline. With each scene written, I’d cross it off and move onto the next. It meant all I really needed to think about was what was directly ahead, not the entire maw that is the rest of the novel. With this book, the further the outline got into the story, the looser it described the events. That hurt me a lot. The less detail I determined early, the more work I had on the day.

New Rules

For me, seven months is too long to write a draft. The longer it takes to write, the more complications crop up along the way. My dream is to draft in 3-4 months. Less than that isn’t possible unless I start increasing my daily word count goals, and I’d rather consistently hit the daily goals I have now than risk pushing myself too hard and lose months from burnout. So, with all that in mind, I’ve set myself a few new rules:

  1. From the moment I start my draft, the next three months can have no major trips, just the occasional weekend getaway.
  2. If I miss 1 day of writing, I have to do everything in my power to make sure I hit my goal the next day.
  3. Once a scene is imagined, it doesn’t get revisited until the day I write it. No over-engineering here.
  4. Outline early, and outline thoroughly.

Hopefully in the near future I’ll be hitting my goal of 2 books a year.

DEBUT NOVEL NOW AVAILABLE FOR PRE-ORDER! (Not the story described in this article):

The Human Countermove is now available for pre-order! https://www.amazon.com/dp/B0FM9R7T5F

In a nation ruled by AI Minds, productivity is everything—even play.

Once a legend in the world of strategy games, Zouk Solinsen is now just another burnout in a society obsessed with efficiency. But when the Minds announce a high-stakes tournament—with a seat on the ruling council as the prize—Zouk is drawn back into the fray, determined to reshape the future.

With help from the enigmatic Torrez Institute, Zouk racks up early victories against the Minds. But when Maya Torrez reveals the cost of her support—a violent coup against the Minds—he rejects it and strikes out alone.

Now, with no allies, dwindling resources, and a nation on the brink, Zouk faces the biggest game of his life—and a final, impossible choice: reform the system from within, or burn it all down.

Some Python Code Proofed My Book in 5 minutes

I wrote my book word by word, no AI involved. An editor helped me develop the story and a copy-editor made sure the manuscript was clean. I’ve read my book about a dozen times. Then my layout person gave me the final version of the book and I realized I had to read the whole thing again to check for new errors.

First I did it properly. My eyes were basically blind by the end. But I wanted a second sweep. The thing is, any person asked to do the job will make a mistake. They’ll overlook something. They won’t realize one paragraph is copied over twice or accidentally cut a space between two sentences. What I needed was a perfect sweep. A complete comparison between my original manuscript and the final epub document. The kind of sweep that could only be performed by a soulless machine with an inflexible view of correct and incorrect.

When I’m not writing I’m coding, and this kind of repetitive, detail-oriented, clearly defined task is the perfect fit for a machine. In fact, it was such a perfect fit, the whole process only took an hour.

What Did The Machine Do?

First I defined my requirements. This code was written to spot exactly one type of problem, copy-and-paste mistakes performed by the layout person. It’s not going to spot typos, it’s not going to spot grammar issues, and it’s certainly not going to point out plot holes. This machine is very stupid, but it performs its job to the letter.

Manuscript format: DOCX

Final Book Layout format: EPUB

Goal: Review every sentence in the EPUB and DOCX files and identify any sentence missing from one file that is present in the other, this should capture any omissions, insertions, or errors in the final manuscript. Then, identify if any sentences appear in the same manuscript more than once, this should identify any ‘duplicate chapter’ or ‘duplicate paragraph’ problems.

The complete code will be shown at the end in case you want to use it, but first I’ll walk you through the parts.

Step 1: Parse the DOCX Manuscript

import docx

def extract_text_from_docx(docx_path):
    doc = docx.Document(docx_path)
    full_text = []
    for para in doc.paragraphs:
        if para.text.strip():  # skip empty paragraphs
            full_text.append(para.text.strip())
    return '\n'.join(full_text)

This code is pretty straightforward, it parses the .docx file into paragraphs, joins it all together into one big paragraphless blob.

Step 2: Parse the EPUB Book

This code is almost identical to the DOCX, but EPUB has a lot more nuance to its data-types. We have to ensure we only retrieve the actual text items, and parse them out of html into plain-text. Then we join it all together in one big wall of book.

import ebooklib
from ebooklib import epub
from bs4 import BeautifulSoup

def extract_text_from_epub(epub_path):
    book = epub.read_epub(epub_path)
    text_content = []

    for item in book.get_items():
        if item.get_type() == ebooklib.ITEM_DOCUMENT:
            soup = BeautifulSoup(item.get_content(), 'html.parser')
            # Remove scripts and styles
            for tag in soup(['script', 'style']):
                tag.decompose()
            text = soup.get_text(separator=' ', strip=True)
            if text:
                text_content.append(text)

    return '\n'.join(text_content)

Step 3: Split the book-blobs into sentences

This part uses a tool called the Natural Language Toolkit (NLTK). Sometimes what NLTK considers a sentence is a little funny, like it’ll join two sets of short quotes together. But we cannot allow perfect to be the enemy of good, so as long as NLTK is responsible for both sentence splitting procedures, the final outputs should be identical.

import nltk
from nltk.tokenize import sent_tokenize
nltk.download('punkt')
nltk.download('punkt_tab')

def split_text_into_sentences(text):
    return sent_tokenize(text)

Step 3: Data cleanup

You may have noticed some really long character replacement stuff. Turns out the docx parser picks up a few too many newlines and the epub parser likes directional quotes, so all of that gets replaced with nice, consistent sentencing.

def docx_scan():
    docx_path = "FILENAME.docx"
    text = extract_text_from_docx(docx_path)
    sentences: List[str] = split_text_into_sentences(text)
    
    for i, val in enumerate(sentences):
        sentences[i] = val.replace('\n', ' ').replace('“', '"').replace('”', '"').replace("‘", "'").replace("’", "'").replace("\'", "'")

    return sentences

def epub_scan():
    epub_path = 'FILENAME.epub'
    text = extract_text_from_epub(epub_path)
    sentences: List[str] = split_text_into_sentences(text)

    for i, val in enumerate(sentences):
        sentences[i] = val.replace('\n', ' ').replace('“', '"').replace('”', '"').replace("‘", "'").replace("’", "'").replace("\'", "'")

    return sentences

Step 4: Crawl through the two books

This is a bit of a doozy, but this function essentially crawls through the final book looking for the next sentence in the manuscript. If it doesn’t find it in 10 sentences, it reports the sentence missing and moves on.
Note: The original draft of this post had a different algorithm that failed to account for sentence order. There’s nothing a programmer does more than tinker with their code, but this function is a big improvement on the original, trust me.

def compare_books(manuscript: List[str], final_book: List[str]):
    # We sweep through final_book searching for sentences from manuscrpt
    book_1_pos: int = 0
    book_2_pos: int = 0
    while book_1_pos < len(manuscript):
        found: bool = False
        target_sentence: str = manuscript[book_1_pos]
        for sweep_position in range(book_2_pos, book_2_pos+10):
            if(sweep_position < len(final_book) and target_sentence == final_book[sweep_position]):
                book_1_pos += 1
                book_2_pos = sweep_position
                found = True
                continue
        
        if not found:
            book_1_pos += 1
            if ' - ' not in target_sentence:
                print(target_sentence)
    

And because of the way the function is written, we can actually crawl through both books the same way.

    epub_sentences = epub_scan()
    docx_sentences = docx_scan()
    # Check the epub file for errors
    compare_books(docx_sentences, epub_sentences)
    # Check the docx file for errors
    compare_books(epub_sentences, docx_sentences)

There are ~8000 sentences in my book. Since the computer reads both copies twice, it’s only about 32,000 operations. A very cheap, less than one second scan for errors.

All the differences are then written out to a file. There were a bunch of false positives. Of the 54 reported omissions, 4 sentences turned out to contain errors, the rest were quirks of the epub format. But finding real errors means it’s working! And it means my layout person did a fantastic job!

Step 5: Check for duplicates

Finally, we do a quick check in both sentence lists for duplicates. The results here reveal my laziness as an author. It turns out I have ~90 non-unique sentences in my book. Most are ‘He said’, ‘She said’, ‘He nodded’, but the strangest one was “Alpha, Golf, Delta, Charlie.” which is a list of squadrons that are referenced in that exact order on two different occasions.

    non_unique_docx = set([x for x in docx_sentences if docx_sentences.count(x) > 1])

    non_unique_epub = set([x for x in epub_sentences if epub_sentences.count(x) > 1])

    print(f"Docx copies: {len(non_unique_docx)}")
    print(f"Epub copies: {len(non_unique_epub)}")

I verified that the total number of non-unique sentences was identical in the DOCX and EPUB formats and moved on.

Conclusion

I always felt a little uneasy about the final version of my book. Even when I had been through it myself, I couldn’t be sure I hadn’t overlooked a massive error. I still can’t be completely sure, but there’s something really reassuring about having a machine do a run-through. When precision is the aim, somehow the passionless report of a calculator is more comforting than a thumbs-up from a professional.

Complete File:

import docx
import nltk
from nltk.tokenize import sent_tokenize
import ebooklib
from ebooklib import epub
from bs4 import BeautifulSoup
from typing import List, Set

nltk.download('punkt')
nltk.download('punkt_tab')

def extract_text_from_docx(docx_path):
    doc = docx.Document(docx_path)
    full_text = []
    for para in doc.paragraphs:
        if para.text.strip():  # skip empty paragraphs
            full_text.append(para.text.strip())
    return '\n'.join(full_text)

def split_text_into_sentences(text):
    return sent_tokenize(text)

def docx_scan():
    docx_path = "YOURFILE.docx"
    text = extract_text_from_docx(docx_path)
    sentences: List[str] = split_text_into_sentences(text)
    
    for i, val in enumerate(sentences):
        sentences[i] = val.replace('\n', ' ').replace('“', '"').replace('”', '"').replace("‘", "'").replace("’", "'").replace("\'", "'")

    return sentences


def extract_text_from_epub(epub_path):
    book = epub.read_epub(epub_path)
    text_content = []

    for item in book.get_items():
        if item.get_type() == ebooklib.ITEM_DOCUMENT:
            soup = BeautifulSoup(item.get_content(), 'html.parser')
            # Remove scripts and styles
            for tag in soup(['script', 'style']):
                tag.decompose()
            text = soup.get_text(separator=' ', strip=True)
            if text:
                text_content.append(text)

    return '\n'.join(text_content)

def epub_scan():
    epub_path = 'YOURFILE.epub'
    text = extract_text_from_epub(epub_path)
    sentences: List[str] = split_text_into_sentences(text)

    for i, val in enumerate(sentences):
        sentences[i] = val.replace('\n', ' ').replace('“', '"').replace('”', '"').replace("‘", "'").replace("’", "'").replace("\'", "'")

    return sentences

def compare_books(manuscript: List[str], final_book: List[str]):
    # We sweep through final_book searching for sentences from manuscript
    book_1_pos: int = 0
    book_2_pos: int = 0
    while book_1_pos < len(manuscript):
        found: bool = False
        target_sentence: str = manuscript[book_1_pos]
        for sweep_position in range(book_2_pos, book_2_pos+10):
            if(sweep_position < len(final_book) and target_sentence == final_book[sweep_position]):
                book_1_pos += 1
                book_2_pos = sweep_position
                found = True
                continue
        
        if not found:
            book_1_pos += 1
            if ' - ' not in target_sentence:
                print(target_sentence)

def main():
    epub_sentences = epub_scan()
    docx_sentences = docx_scan()
    compare_books(docx_sentences, epub_sentences)
    compare_books(epub_sentences, docx_sentences)

    non_unique_docx = set([x for x in docx_sentences if docx_sentences.count(x) > 1])
    non_unique_epub = set([x for x in epub_sentences if epub_sentences.count(x) > 1])
    print(f"Docx copies: {len(non_unique_docx)}")
    print(f"Epub copies: {len(non_unique_epub)}")


if __name__ == '__main__':
    main()