What Pixelhop learned when Glitch missed a known-good business and configured an Instantly follow-up incorrectly.
Glitch is our Hermes Agent, a persistent AI agent we run across Pixelhop. Unlike a standalone chat, Glitch can read and update our company knowledge base, research businesses across the web, and use connected tools such as Instantly to configure and check campaigns. It can also carry out multi-step tasks and coordinate other AI models such as Claude.
Finding the right businesses is exactly the kind of job I wanted to try it on. There is a lot of repetitive searching involved, but the useful part still comes down to knowing what a good result looks like. We wanted to see how much of the grunt work Glitch could take on while we kept hold of that judgement.
At first glance, the list looked promising. It had plenty of companies on it and enough detail to feel useful, but it missed Rye & Beyond, a holiday cottage business I had deliberately given Glitch as an example of the kind of company it should find. Once I saw that, the size and polish of the rest of the list mattered a lot less. That one omission told us more than everything else Glitch had found.
We kept refining the search until it could recover the example, then got much more precise about what made a business a good fit. That part started to work much better. Then Glitch made a completely different mistake while setting up our Instantly sequence. The follow-up looked as though it was scheduled four days after the first email, but the admin screen had stored a zero-day delay and sent it about half an hour later, which is not the introduction you want to make to a business you have carefully chosen.
Neither problem was disastrous, but both were useful. They reminded me how easy it is to trust something because it looks complete, whether that is a tidy list with lots of detail or a message from Glitch saying the setup is done.

Figure: How we split the research and verification work between Gem, Glitch and Instantly.
These are the five checks we are now using.
1. Start with a company Glitch must find
Before asking Glitch to search widely, I gave it a business I already knew should qualify. Rye & Beyond was not a lucky example we found afterwards. It was a deliberate test of the search.
When the first pass missed it, we knew the process was wrong even though the rest of the list looked convincing. Without that check, we might have spent time reviewing a large set of results built on a weak definition.
This has become our first rule for AI-assisted research: begin with one or two known-good examples. It sounds obvious written down, but it was not how we started. If the search cannot recover them, do not make the list bigger. Work out what the examples reveal about the target and change the search until it can pass that test.

Figure: The supplied company gave us a pass-or-fail check before we trusted the wider list.
2. Define a good lead before asking for more leads
Our early request was too broad. We wanted independent holiday cottage businesses, but that can include directories, hotels, national operators, residential agents and single-property hosts. A search can return all of those and still look busy and productive.
We replaced the broad description with explicit inclusion and exclusion rules. We wanted businesses with their own website, direct booking or enquiry, more than one property or a professionally managed portfolio, and enough operational complexity for Chat Thing to be relevant. We excluded directories, online travel agencies, hotels, residential agents and businesses that were structurally very different from our known-good example. Writing those rules down was where the useful work happened, because Glitch was never short of companies to suggest. The difficult bit was showing it what we actually meant by the right company.
Once we had done that judgement work, Glitch became much more useful. We did not need a cleverer prompt as much as we needed a clearer shared definition of what belonged on the list and what did not.
3. Give each tool a separate job
We started by treating business discovery, contact finding and sending as one long job. In practice, asking one tool to be good at all of it made the whole process harder to trust, so we split it up.
For this niche, browser-led research with Glitch was more useful for finding smaller independent businesses and reading what they actually did. Instantly and Apollo appeared stronger later in the process, when we needed to find people, verify work email addresses and manage controlled sending. That is an observation from these searches rather than a universal verdict on either product.
I still made the final qualification decision. Glitch could research, compare and organise candidates, but a company only moved forward when it looked like a business we had a real reason to contact.
4. Keep the first cohort small
We are testing this across two groups: B Corps for Pixelhop and holiday cottage businesses for Chat Thing. We could have taken every company Glitch found and dropped them all into an outreach campaign, but a bigger send would only have created more activity while we were still working out whether the process was any good.
We reviewed the candidates, chose small initial cohorts and kept automatic replies, human replies, bounces, opportunities and customers separate. At our 30 August checkpoint, the pilots had contacted 21 businesses and had not received a human reply. That means we can talk about the process we are testing, but not claim that it generates leads or customers.
The small cohorts give us room to spot a poor fit, an awkward message or a broken sequence before the mistake reaches a much larger list. That matters more to me than making the numbers look busy. We are trying to learn whether the workflow can produce worthwhile conversations while the method is still changing, and sending more emails would not make that answer arrive any faster.
5. Check the real system after Glitch says it is done
Glitch filled out the Instantly sequence incorrectly. The API representation appeared to show a four-day delay, while the admin screen held a zero-day delay. The first follow-up went out about 31 minutes after the opening email.
Instantly sent the follow-up according to the stored sequence. Glitch had configured that sequence incorrectly, and the difference between the machine-readable representation and the human-visible setup made the mistake harder to spot.
We corrected both campaigns and added a simple read-back step. After changing a sequence, we reload the admin screen, read every delay as a person will see it and test the timing with one lead before increasing the volume. It is not an exciting extra step, but neither is explaining why somebody received a follow-up 31 minutes after the first email. A successful tool call or completion message tells us that something ran; we still need to look at the real system and decide whether it did what we intended.

Figure: The corrected Instantly sequence. We now reopen the real campaign and read each saved delay before increasing volume.
Where we are now
This is still an early experiment, and I do not want to turn a better process into a success story before it has earned one. We have improved the way Glitch finds and filters businesses, separated discovery from contact verification and sending, and added a real-world check after browser-led setup. We have not yet proved that the process produces replies, opportunities or customers.
The useful part so far is the method: start with a known-good example, define fit before chasing volume, give each tool a bounded job, keep the first cohort small and verify the human-visible result. I still want Glitch doing this kind of work because there is no prize for spending my afternoon opening dozens of websites. What this test has made clearer is where I want to stay involved: choosing the examples, deciding what a good business looks like and checking the real result before we contact more people. Glitch can do plenty of the grunt work, but I do not want to outsource the judgement that tells us whether any of it was worth doing.