Module 2

Building Your Lead Research Agent

Automate prospect discovery with AI agents, web scraping, and smart databases.

45
Minutes
15
Slides
50+
Leads/Day Output
1
Working Agent
1 / 15
Module 2

The Lead Gen Pipeline — In Depth

Module 1 introduced the 6 stages. Now let's zoom into the research stage — where your agent will operate.

01

Research

Find ideal prospects

02

Qualify

Score & enrich

03

Outreach

Personalized messages

04

Convert

Chatbots & follow-up

05

Nurture

Build relationships

06

Optimize

Measure & improve

🔍

Why Research Matters

Research is the foundation. Bad data in = bad results out. Your agent ensures every lead is accurate, relevant, and enriched.

⚡

The Research Stage

Your agent will: discover prospects → scrape data → deduplicate → enrich → score → output to database. Fully automated.

2 / 15
Module 2

What is a Lead Research Agent?

A Lead Research Agent is an AI-powered system that automatically discovers, collects, and profiles potential customers.

🤖

Core Capabilities

  • • Discovers prospects from multiple sources
  • • Scrapes contact info & firmographics
  • • Enriches data with AI analysis
  • • Deduplicates and cleans records
  • • Scores leads by fit
  • • Outputs to your CRM/database
🧠

How It Works

The agent uses a combination of:

  • • LLMs for understanding & extraction
  • • Web scraping for data collection
  • • APIs for enrichment
  • • Databases for storage
  • • Automation for scheduling

💡 Key Insight

A lead research agent isn't just a scraper — it's an intelligent system that understands your ICP and finds prospects that match it.

3 / 15
Module 2

Setting Up Hermes Agent for Lead Research

Hermes Agent is the AI engine that powers your lead research. Let's get it configured.

📦

Installation

$ pip install hermes-agent
Successfully installed hermes-agent-2.1.0
$ hermes init
✓ Configuration file created
✓ Default profile initialized
✓ Ready to use
⚙️

Configuration

# ~/.hermes/config.yaml
model: "gpt-4o"
temperature: 0.7
max_tokens: 4096

tools:
  - web_search
  - web_extract
  - browser
  - terminal

lead_research:
  sources:
    - linkedin
    - company_websites
    - directories
  output_format: "google_sheets"
  schedule: "daily_9am"
4 / 15
Module 2

Creating Your First Lead Research Agent

Time to build! Here's a complete lead research agent in Python.

import hermes
from hermes.tools import web_search, web_extract, sheets

class LeadResearchAgent:
    def __init__(self, icp):
        self.icp = icp  # Ideal Customer Profile
        self.leads = []

    def research(self, query, max_results=50):
        """Search for prospects matching ICP"""
        results = web_search(
            query=self._build_query(query),
            limit=max_results
        )
        return self._parse_results(results)

    def _build_query(self, base):
        return (
            f"{base} "
            f"{self.icp['industry']} "
            f"{self.icp['title']} "
            f"{self.icp['location']}"
        )

    def _parse_results(self, results):
        for r in results:
            lead = {
                'name': r.get('title'),
                'company': r.get('site'),
                'url': r.get('url'),
                'source': 'web_search'
            }
            self.leads.append(lead)
        return self.leads

# Usage
agent = LeadResearchAgent(icp={
    'industry': 'SaaS',
    'title': 'CTO OR VP Engineering',
    'location': 'United States'
})
leads = agent.research('engineering leadership')
sheets.append('Leads', leads)
5 / 15
Module 2

Web Scraping for Prospects

Python + BeautifulSoup is the most effective way to extract structured lead data from websites.

import requests
from bs4 import BeautifulSoup
import re

def scrape_prospects(url):
    """Scrape company directory for leads"""
    headers = {
        'User-Agent': 'LeadResearchBot/1.0'
    }
    response = requests.get(url, headers=headers)
    soup = BeautifulSoup(response.text, 'html.parser')

    leads = []
    # Find all company cards
    for card in soup.select('.company-card'):
        lead = {
            'company': card.select_one('.company-name').text.strip(),
            'website': card.select_one('a')['href'],
            'industry': card.select_one('.industry').text.strip(),
            'size': card.select_one('.size').text.strip(),
            'location': card.select_one('.location').text.strip(),
        }
        leads.append(lead)

    return leads

# Scrape multiple pages
for page in range(1, 11):
    url = f"https://directory.example.com?page={page}"
    leads = scrape_prospects(url)
    # Save to database...

🎯 Best Practices

  • • Respect robots.txt
  • • Add delays between requests
  • • Use proper User-Agent
  • • Handle pagination
  • • Cache results locally

⚠️ Common Pitfalls

  • • Scraping too fast (get blocked)
  • • Not handling errors
  • • Ignoring JavaScript-rendered content
  • • Forgetting to deduplicate
  • • Not validating email formats
6 / 15
Module 2

Building Lead Databases with Google Sheets

Google Sheets is a free, flexible CRM for storing and managing your leads. Here's the schema we'll use.

📊 Lead Database Schema

Column Type Description Example
lead_id String Unique identifier lead_001
name String Contact name Jane Smith
email String Email address jane@company.com
company String Company name Acme Inc
title String Job title VP Engineering
industry String Industry vertical SaaS
company_size String Employee count 50-200
location String Geographic location San Francisco, CA
source String Lead source web_scrape
score Number Lead score (0-100) 85
status String Pipeline status new
7 / 15
Module 2

Lead Deduplication & Cleaning

Raw scraped data is messy. Here's how to clean and deduplicate your leads automatically.

import re
from difflib import SequenceMatcher

def clean_leads(leads):
    """Clean and deduplicate lead list"""
    cleaned = []
    seen = set()

    for lead in leads:
        # Normalize email
        email = lead['email'].lower().strip()
        if not re.match(r'^[\w\.-]+@[\w\.-]+\.\w+$', email):
            continue  # Skip invalid emails

        # Create dedup key
        key = email.split('@')[0] + '@' + email.split('@')[1].split('.')[0]

        if key in seen:
            continue
        seen.add(key)

        # Clean fields
        lead['name'] = lead['name'].title().strip()
        lead['company'] = lead['company'].strip()
        lead['email'] = email
        cleaned.append(lead)

    return cleaned

def fuzzy_dedup(leads, threshold=0.85):
    """Remove near-duplicate company names"""
    unique = []
    for lead in leads:
        is_dup = False
        for existing in unique:
            ratio = SequenceMatcher(
                None,
                lead['company'].lower(),
                existing['company'].lower()
            ).ratio()
            if ratio > threshold:
                is_dup = True
                break
        if not is_dup:
            unique.append(lead)
    return unique
8 / 15
Module 2

Automating the Research Workflow

Connect all the pieces into a fully automated pipeline that runs on schedule.

🕐 Trigger

Schedule (cron)

→

🔍 Discover

Web search

→

🕷️ Scrape

BeautifulSoup

→

🧹 Clean

Deduplicate

→

📊 Store

Google Sheets

# automation.py — Run daily at 9 AM
import schedule
import time
from lead_agent import LeadResearchAgent
from scraper import scrape_prospects
from cleaner import clean_leads, fuzzy_dedup

def daily_research():
    """Full automated research pipeline"""
    agent = LeadResearchAgent(icp=ICP)

    # 1. Discover new prospects
    new_leads = agent.research('target keywords')

    # 2. Scrape additional data
    scraped = []
    for lead in new_leads:
        data = scrape_prospects(lead['url'])
        scraped.extend(data)

    # 3. Clean & deduplicate
    cleaned = clean_leads(scraped)
    unique = fuzzy_dedup(cleaned)

    # 4. Store in database
    sheets.append('Leads', unique)
    print(f"Added {len(unique)} new leads")

# Schedule daily at 9:00 AM
schedule.every().day.at("09:00").do(daily_research)

while True:
    schedule.run_pending()
    time.sleep(60)
9 / 15
Module 2

Scaling Your Research Agent

Start small, then scale to thousands of leads per day with these strategies.

📈

Volume Scaling

  • • Multiple search queries
  • • Parallel scraping
  • • Proxy rotation
  • • Distributed workers
  • • Rate limit management
🎯

Quality Scaling

  • • AI-powered scoring
  • • ICP refinement
  • • Source prioritization
  • • Enrichment APIs
  • • Human-in-the-loop review
🔧

Infrastructure

  • • Cloud hosting (AWS/GCP)
  • • Database upgrade (PostgreSQL)
  • • Queue system (Redis)
  • • Monitoring & alerts
  • • Backup & recovery
100
Leads/Day (Solo)
1,000
Leads/Day (Team)
10,000
Leads/Day (Scale)
10 / 15
Module 2

Module 2 Recap

Let's review what we've covered:

📚 Key Concepts

  • • Lead Research Agent = AI + scraping + database
  • • Hermes Agent powers the intelligence layer
  • • BeautifulSoup extracts structured data
  • • Google Sheets = free, flexible CRM
  • • Deduplication is critical for data quality
  • • Automation runs the pipeline on schedule

🛠️ What You Built

  • • Configured Hermes Agent for lead research
  • • Created a LeadResearchAgent class
  • • Built web scraper with BeautifulSoup
  • • Designed Google Sheets lead database
  • • Implemented deduplication & cleaning
  • • Automated the full research workflow
11 / 15
Module 2

Module 2 Quiz

Test your understanding before moving to Module 3.

Question 1

What are the 4 core components of a Lead Research Agent?

Answer: LLM (intelligence), web scraping (collection), APIs (enrichment), database (storage)

Question 2

Why is deduplication important in lead generation?

Answer: Duplicate leads waste outreach budget, skew metrics, and create poor customer experiences. Clean data = better results.

Question 3

Name 3 ways to scale a lead research agent beyond 100 leads/day.

Answer: Parallel scraping, multiple search queries, proxy rotation, distributed workers, cloud infrastructure (any 3).

12 / 15
Module 2

What's Next: Module 3

In Module 3, you'll take your leads and turn them into qualified opportunities with AI scoring and enrichment.

🚀 Module 3 Preview: AI Lead Scoring & Enrichment

  • • Building AI lead scoring models
  • • Enriching leads with firmographic data
  • • Predicting conversion probability
  • • Integrating with CRM systems
  • • Output: A prioritized list of sales-ready leads
13 / 15
Module 2

Your Next Steps

Before starting Module 3, make sure you've completed these action items:

✅ Action Items

  • Set up Hermes Agent on your machine
  • Create your Google Sheets lead database
  • Build and test your first scraper
  • Run your first automated research cycle
  • Collect at least 50 leads

📝 Notes

  • • Save all code to your project folder
  • • Document your ICP for reference
  • • Test with small batches first
  • • Join the course community for help
  • • Review Module 1 if needed
14 / 15
Module 2

Module 2 Complete! 🎉

You've built a working Lead Research Agent that can find and profile prospects automatically.

Ready for Module 3?

Let's add AI scoring and enrichment to turn your raw leads into sales-ready opportunities.

15 / 15