Automate prospect discovery with AI agents, web scraping, and smart databases.
Module 1 introduced the 6 stages. Now let's zoom into the research stage — where your agent will operate.
Find ideal prospects
Score & enrich
Personalized messages
Chatbots & follow-up
Build relationships
Measure & improve
Research is the foundation. Bad data in = bad results out. Your agent ensures every lead is accurate, relevant, and enriched.
Your agent will: discover prospects → scrape data → deduplicate → enrich → score → output to database. Fully automated.
A Lead Research Agent is an AI-powered system that automatically discovers, collects, and profiles potential customers.
The agent uses a combination of:
A lead research agent isn't just a scraper — it's an intelligent system that understands your ICP and finds prospects that match it.
Hermes Agent is the AI engine that powers your lead research. Let's get it configured.
# ~/.hermes/config.yaml
model: "gpt-4o"
temperature: 0.7
max_tokens: 4096
tools:
- web_search
- web_extract
- browser
- terminal
lead_research:
sources:
- linkedin
- company_websites
- directories
output_format: "google_sheets"
schedule: "daily_9am"
Time to build! Here's a complete lead research agent in Python.
import hermes
from hermes.tools import web_search, web_extract, sheets
class LeadResearchAgent:
def __init__(self, icp):
self.icp = icp # Ideal Customer Profile
self.leads = []
def research(self, query, max_results=50):
"""Search for prospects matching ICP"""
results = web_search(
query=self._build_query(query),
limit=max_results
)
return self._parse_results(results)
def _build_query(self, base):
return (
f"{base} "
f"{self.icp['industry']} "
f"{self.icp['title']} "
f"{self.icp['location']}"
)
def _parse_results(self, results):
for r in results:
lead = {
'name': r.get('title'),
'company': r.get('site'),
'url': r.get('url'),
'source': 'web_search'
}
self.leads.append(lead)
return self.leads
# Usage
agent = LeadResearchAgent(icp={
'industry': 'SaaS',
'title': 'CTO OR VP Engineering',
'location': 'United States'
})
leads = agent.research('engineering leadership')
sheets.append('Leads', leads)
Python + BeautifulSoup is the most effective way to extract structured lead data from websites.
import requests
from bs4 import BeautifulSoup
import re
def scrape_prospects(url):
"""Scrape company directory for leads"""
headers = {
'User-Agent': 'LeadResearchBot/1.0'
}
response = requests.get(url, headers=headers)
soup = BeautifulSoup(response.text, 'html.parser')
leads = []
# Find all company cards
for card in soup.select('.company-card'):
lead = {
'company': card.select_one('.company-name').text.strip(),
'website': card.select_one('a')['href'],
'industry': card.select_one('.industry').text.strip(),
'size': card.select_one('.size').text.strip(),
'location': card.select_one('.location').text.strip(),
}
leads.append(lead)
return leads
# Scrape multiple pages
for page in range(1, 11):
url = f"https://directory.example.com?page={page}"
leads = scrape_prospects(url)
# Save to database...
Google Sheets is a free, flexible CRM for storing and managing your leads. Here's the schema we'll use.
| Column | Type | Description | Example |
|---|---|---|---|
| lead_id | String | Unique identifier | lead_001 |
| name | String | Contact name | Jane Smith |
| String | Email address | jane@company.com | |
| company | String | Company name | Acme Inc |
| title | String | Job title | VP Engineering |
| industry | String | Industry vertical | SaaS |
| company_size | String | Employee count | 50-200 |
| location | String | Geographic location | San Francisco, CA |
| source | String | Lead source | web_scrape |
| score | Number | Lead score (0-100) | 85 |
| status | String | Pipeline status | new |
Raw scraped data is messy. Here's how to clean and deduplicate your leads automatically.
import re
from difflib import SequenceMatcher
def clean_leads(leads):
"""Clean and deduplicate lead list"""
cleaned = []
seen = set()
for lead in leads:
# Normalize email
email = lead['email'].lower().strip()
if not re.match(r'^[\w\.-]+@[\w\.-]+\.\w+$', email):
continue # Skip invalid emails
# Create dedup key
key = email.split('@')[0] + '@' + email.split('@')[1].split('.')[0]
if key in seen:
continue
seen.add(key)
# Clean fields
lead['name'] = lead['name'].title().strip()
lead['company'] = lead['company'].strip()
lead['email'] = email
cleaned.append(lead)
return cleaned
def fuzzy_dedup(leads, threshold=0.85):
"""Remove near-duplicate company names"""
unique = []
for lead in leads:
is_dup = False
for existing in unique:
ratio = SequenceMatcher(
None,
lead['company'].lower(),
existing['company'].lower()
).ratio()
if ratio > threshold:
is_dup = True
break
if not is_dup:
unique.append(lead)
return unique
Connect all the pieces into a fully automated pipeline that runs on schedule.
Schedule (cron)
Web search
BeautifulSoup
Deduplicate
Google Sheets
# automation.py — Run daily at 9 AM
import schedule
import time
from lead_agent import LeadResearchAgent
from scraper import scrape_prospects
from cleaner import clean_leads, fuzzy_dedup
def daily_research():
"""Full automated research pipeline"""
agent = LeadResearchAgent(icp=ICP)
# 1. Discover new prospects
new_leads = agent.research('target keywords')
# 2. Scrape additional data
scraped = []
for lead in new_leads:
data = scrape_prospects(lead['url'])
scraped.extend(data)
# 3. Clean & deduplicate
cleaned = clean_leads(scraped)
unique = fuzzy_dedup(cleaned)
# 4. Store in database
sheets.append('Leads', unique)
print(f"Added {len(unique)} new leads")
# Schedule daily at 9:00 AM
schedule.every().day.at("09:00").do(daily_research)
while True:
schedule.run_pending()
time.sleep(60)
Start small, then scale to thousands of leads per day with these strategies.
Let's review what we've covered:
Test your understanding before moving to Module 3.
What are the 4 core components of a Lead Research Agent?
Answer: LLM (intelligence), web scraping (collection), APIs (enrichment), database (storage)
Why is deduplication important in lead generation?
Answer: Duplicate leads waste outreach budget, skew metrics, and create poor customer experiences. Clean data = better results.
Name 3 ways to scale a lead research agent beyond 100 leads/day.
Answer: Parallel scraping, multiple search queries, proxy rotation, distributed workers, cloud infrastructure (any 3).
In Module 3, you'll take your leads and turn them into qualified opportunities with AI scoring and enrichment.
Before starting Module 3, make sure you've completed these action items:
You've built a working Lead Research Agent that can find and profile prospects automatically.
Let's add AI scoring and enrichment to turn your raw leads into sales-ready opportunities.