Skip to main content
Jimmy Jansen

A Pragmatic Testing Strategy

Coverage numbers are easy to hit and easy to trust too much. A useful testing strategy makes the important failures visible and gives the team feedback it can act on.

I want to be able to refactor a difficult piece of code and trust I haven’t broken something three modules away. When someone newer to the codebase makes a change, I want the suite to catch the mistake I would have caught by hand. A green build is useful when I can explain which of those worries it has checked.

I used to dictate testing strategy. Now I mostly try to guide teams toward what works for their context. I’ve had this conversation with developers, product owners, management, and leadership, and gotten plenty of it wrong along the way. I’ve also seen teams succeed with very different approaches. The choices below are the ones I keep coming back to.

Link to section: Start with the failureStart with the failure

Deciding what to test is most of the skill. I start with the consequence of getting it wrong, how likely I am to make that mistake, and whether the test would catch it. A permission check can be one line long and still decide whether someone gets access to another customer’s data. Its size tells me very little about how much protection it deserves.

I also test the things I don’t expect to get right the first time. I’m confident I can write a sliding window, but it is fiddly enough that I’ll write tests around it. Boundary cases are where my own mistakes live. That judgment gets easier with experience, though familiarity is a poor excuse to leave an important rule unchecked.

Coverage can help me find code the suite never exercises. It can’t tell me whether the assertions would notice the failure I care about. I’ve seen tests check getters and setters, confirm that mocks returned their configured values, and rubber-stamp snapshots while the product still broke. Before adding another test to move the number, I want to know what mistake it would expose.

Difficulty writing that test can be useful feedback too. If I need a database just to exercise a calculation, or fifteen mocks to reach one decision, I look at how the code is put together. Perhaps business logic and infrastructure are tangled, or the function has too many jobs. Sometimes the work really is awkward to test. I treat it as a reason to investigate the design, rather than proof that the design is wrong.

Link to section: Make a difficult system safer to changeMake a difficult system safer to change

One of my assignments at Rangle involved a codebase north of ten million lines with essentially no tests and recurring daily outages lasting hours. Small changes could ripple through undocumented behaviour. The immediate job was to stop those outages, and we needed a way to tell whether our changes had broken something else.

We established a production-like development environment, measurements, and black-box tests around critical workflows. Those tests exercised the system from the outside, giving us checks we could repeat while changing the internals. Trying to cover the whole codebase would have delayed the urgent work. We needed enough protection around the paths that had to keep working.

That gave us a way to check the database upgrades and query changes as the remediation progressed. By the end of the eight-week engagement, we had resolved the recurring daily outages. The project was then halted and I was taken off it.

Link to section: Choose how much of the system to testChoose how much of the system to test

Once I know the failure I’m trying to catch, I decide what needs to be real in the test. I think of this as turning a dial: let more of the real system in, and check more of the assumptions I would otherwise have to simulate.

That usually adds setup and upkeep, but it isn’t a strict ranking of speed or reliability. A focused test against a real database can be easier to understand than a broad test with a pile of mocks. Scope, isolation, and how easily I can diagnose a failure matter too.

Teams draw the boundaries between test types differently. These are the labels I use.

Link to section: Unit testsUnit tests

I use unit tests for a piece of logic in isolation: a calculation, a business rule, an algorithm. They let me check lots of cases without setting up the rest of the application. I like keeping them running while I work, so a mistake shows up while the code is still in my head.

If a rule decides access or calculates money, I want the expected results to come from that rule. Copying the implementation into the test can reproduce the same mistake on both sides. A small test with an independently chosen expected result can protect something important.

Link to section: Component testsComponent tests

Component tests are the backbone of my suites. I wire several units together and replace the boundaries with mocks or stubs. The code inside the component is real, but the database or external service it talks to is simulated. That keeps the feedback focused on the behaviour I’m changing, without requiring the rest of the environment to be available.

These tests let me check how the pieces work together, including what happens when a dependency fails. Checking that my code makes the right call can be useful. Calling a mock directly and confirming that it returns the value I configured checks nothing about my application.

For years I called these integration tests. Plenty of teams still do. Whatever the label, the limit matters: if the fake behaves differently from the real dependency, a passing component test won’t settle that disagreement.

Link to section: Integration testsIntegration tests

I use integration tests to check a real boundary: the database, a queue, or another service. The point is to exercise the interaction the application depends on. A mock can’t establish that a migration works against the relevant schema or that a query behaves correctly under the database’s constraints.

I keep a smaller set of these checks alongside the component tests. A failure here can expose a wrong assumption in the application, an incompatible dependency, or a problem with the test environment. Keeping each check focused helps me work out which of those I’m dealing with.

Link to section: E2E testsE2E tests

End-to-end tests exercise a user journey through a real interface. Being able to log in and access something you’ve paid for depends on more than the individual rules passing. I want a few tests that follow those critical paths through the system, with a real browser, real application services, and realistic data.

I like using browser automation while building a feature, but I’m selective about which tests I keep in CI. I’ve seen teams try to cover everything through the browser and end up with a suite that takes forever. Debugging it feels like solving a murder mystery where all the witnesses are lying. A failure could be anywhere along the journey.

I keep the important flows in CI and put detailed rule and edge-case coverage closer to the code responsible for them.

Link to section: Run the relevant checks before releaseRun the relevant checks before release

Fast feedback matters because I want to use it throughout the work. Unit and component tests are useful on every change. I usually run the broader set of real-boundary checks on a schedule, often nightly, to look for migration and ORM drift, schemas that no longer match the code, and changes in external contracts.

That schedule doesn’t decide whether a particular change is ready to ship. If a change affects a schema, migration, or service contract, the relevant check against that boundary belongs before release. A green suite against mocked boundaries can’t establish compatibility. Critical E2E flows in CI can cover some of that risk, depending on what they exercise. A login test won’t tell me whether a data migration preserves existing records.

The broader suite can still run on its own schedule. When releases happen during the day, moving every check into a nightly job can let a failure reach customers before the job runs. Running the entire system for every edit can make feedback slow enough that people stop using it.

Link to section: Leave room for QA to surprise youLeave room for QA to surprise you

Every time I was sure my code was bulletproof, a good QA engineer needed about five minutes to prove otherwise. They used the back button. They opened it in two tabs at once. They put an emoji in the name field. In other words, they did what real users do.

My automated tests start from the behaviour and failure cases I thought to describe. A QA engineer brings another way of looking at the product, including interactions I hadn’t considered. That work deserves room in the strategy. A green suite doesn’t make it redundant.

Once we understand a failure, we can decide where a regression test belongs. Finding it through the browser doesn’t mean every future check has to use the browser. If a focused test can catch the same mistake, it can give us faster feedback while QA keeps exploring what we haven’t thought of yet.

Link to section: Start with the team’s problemStart with the team’s problem

I don’t start by asking a team to copy a particular test pyramid. With no tests, I’d pick one important path and get a useful check running. With a flaky suite, I’d fix the failures that have taught people to ignore it. With too much maintenance, I’d look for tests that break when the implementation changes even though the behaviour still works.

Deadline pressure is real. Sometimes the team has to leave something unchecked, and that decision should include the likely cost of debugging it in production, interrupting other work, or recovering from an outage. I’d protect the most consequential behaviour first and make the remaining gap explicit: what is unchecked, who owns it, and when it needs to be closed.

The team needs to understand those choices well enough to maintain them. For the next change, I want us to be able to name the failure we’re trying to prevent, show the check that would catch it, and say what we’re still relying on people to verify before we ship.