Scrapy jobs
We have thousands of websites we need scraped and are looking for someone who has done this professionally. You need to be able to run many scrapers simultaneously and deliver clean data in CSV form...run many scrapers simultaneously and deliver clean data in CSV format. This is a long term position for the right person. We are looking for someone who can dedicate full focus and time to this work rather than splitting attention across many clients. If the fit is right, this is ongoing work for a year or more. Must know Python, Scrapy, Playwright or Selenium. Starting with a paid test task to verify quality before committing. Skills: Web Scraping, Python, Scrapy, Selenium, Data Mining, Data Extraction, Beautiful Soup, CSV How to pay: Pay by the hour Estimated budget: $2 t...
...free to add Craigs list, CareerBuilder, or niche sites as long as they host downloadable resumes or profile data. • Industries: no narrow filter—I’m interested in technology, healthcare, finance, hospitality, construction, and any “various jobs” you come across. • Geography: restrict to the greater Las Vegas area (use location filters or keywords). Technical notes A Python solution using Scrapy, BeautifulSoup, Selenium, or a comparable stack is fine if it avoids detection and respects rate limits. Handle CAPTCHAs or login requirements gracefully. Please build in simple configuration so I can adjust city or keyword filters later without touching code. Deliverables 1. Clean CSV file and an identical Excel workbook containing all scraped r...
I need a fully-functioning lead-generation and sales operating system tailored to my granite and stone business and restricted to prospects located in Pune city. Core re...Export to a clean CSV or direct sync with the CRM I specify. 4. Simple dashboard where I can trigger a new crawl, view statistics and download the latest file. Acceptance criteria • Minimum 95 % valid email and phone coverage on the final list. • All leads mapped to Pune; no out-of-city records. • Source code, documentation and hand-off session included. Python, Selenium, BeautifulSoup, Scrapy or similar tools are fine; feel free to recommend alternatives if they boost reliability. The sooner the first usable dataset is in my hands, the better—I’m ready to move ahead as soon a...
I need a clean, up-to-date scrape of Google local results covering all of Southern California. The focus is on three business types—Hair Extension Salons, Botox Spas, and general Hair Salons. For every location you find, I want the fo...Address, City, Zip • Phone Number • Website URL • Email Address (pull it from the site or listing if shown) Please deliver everything in a single, well-structured CSV file so I can load it straight into my CRM. Accuracy matters more than sheer volume; duplicate removal, consistent formatting, and clear separation of each field are essential acceptance criteria. Feel free to use Python, Scrapy, Selenium, Google Maps API, or any combination of tools you trust—as long as the final dataset is complete, deduplicated,...
I have a specific website from which I need all available email addresses together with the associated names and physical ...need all available email addresses together with the associated names and physical addresses extracted in a single, clean pass. This is a one-time scrape, not an ongoing collection, so accuracy and completeness on the first delivery are critical. Please return the results in a single CSV file, clearly labelled columns for Email, Name, and Address. If you automate with Python, BeautifulSoup, Scrapy, Selenium, or a similar toolset, that’s great—just make sure the final file opens correctly in any spreadsheet editor. Deliverable • CSV containing every accessible email, name, and address from the target site, with no duplicates and valid f...
I already have flat-image mock-ups of every page, so the visual side is locked in; I simply need those screens turned into a clean, responsive informational website. Beyond the front-end build, the key requirement is a small data-puller that...I can trigger manually or via a cron job is fine. Please: • Recreate each page exactly as shown in my images (HTML5/CSS3/JS or a framework you prefer). • Build a simple admin or script where I can run “update” and watch the catalogue refresh. • Keep code well-commented so I can tweak selectors or timing later. If you’re comfortable with typical scraping stacks—Python (BeautifulSoup, Scrapy), Node, or PHP cURL—and can hand over the source along with basic deployment instructions, this ...
...define through a settings panel • A simple sales-trend module so I can see how each item has moved over the past weeks or months and spot consistent winners I’m happy with a cron job, Windows Task Scheduler or any lightweight scheduler you prefer, as long as the update cadence is reliable and I can trigger a manual refresh when needed. Preferred stack is flexible—Python with BeautifulSoup/Scrapy, Node with Puppeteer or any robust combination you’re comfortable maintaining. What matters is clean, well-annotated code, clear setup instructions and resilience to minor site layout changes. Deliverables I expect: – Executable (desktop or headless script) plus source code – Modules that pull data from each of the three marketplaces, map ...
...and all available photos so I can revisit the brochure months after the original ad has vanished. I will pay a set up and small monthly management fee of the process Data format isn’t fixed—CSV, Excel, JSON or anything similarly straightforward all work as long as I can open it later. What matters is that the archive is tidy and searchable. Key deliverables • A dependable scraper (Python, Scrapy/Selenium/BeautifulSoup—your call) scheduled to run weekly without manual input. • De-duping logic based on listing ID or URL so the same property is never saved twice. • For each listing, a folder containing the pictures plus a metadata file holding price and description. • Initial back-fill of current live listings so the archive starts c...
I need a clean, up-to-date extract of specific listings from an accommodation website. The focus is on Apartments, Cottages, and Bed & Breakfast properties, and for each listing I want all st...types covered (apartments, cottages, B&Bs) • Accurate capture of property name, full address, geo-coordinates if shown, amenities summary, and contact phone number • No duplicates, no missing rows, no truncated text • Scraper respects and rate limits or uses rotating proxies where needed • Clear, well-commented code supplied so I can rerun the scrape later (Python, BeautifulSoup/Scrapy/Selenium—use what suits you) If anything on the site requires special handling—JavaScript-rendered content, pagination, or CAPTCHA—please note it up front...
...name • Full address • Total area in acres (or square feet if acres are missing) • Number of houses or flats in the project • Key amenities mentioned by the developer or seller • Cost per house / flat (latest listed price or price range) Deliverables 1. A CSV or Excel file with at least 1,000 complete rows that meet the above criteria. 2. The scraping script (Python + BeautifulSoup/Scrapy, Node.js with Cheerio/Puppeteer, or a comparable stack) well-commented so I can rerun it later. 3. Brief run instructions and any environment requirements. Acceptance criteria • No duplicate projects. • All mandatory fields populated for 95 %+ of rows. • Only Bengaluru addresses; anything outside the city will be rejected. Ple...
...to full notice, etc.) captured consistently across all sites. If additional obvious fields appear, include them as well. Output format Everything must arrive in a single, well-structured Excel spreadsheet, one row per tender with clear column headings. Please normalise dates and strip duplicates so the file is immediately usable for analysis. Technical approach Feel free to employ Python, Scrapy, BeautifulSoup, Selenium or any combination that handles pagination, dynamic content or CAPTCHA challenges. The key is repeatability: I’d like to rerun the script later, so provide clean, documented code. Deliverables • Final Excel workbook containing all scraped tender records • Runnable script or notebook with brief setup instructions • I will want the ...
We are looking for an experien...strictly check existing product SKUs to prevent duplicate entries or overwrite errors. Only new products should be added or existing ones updated safely. Bulk Import: Format the final data into a clean CSV file and successfully import it into WooCommerce without crashing the server (batch imports recommended). Requirements for Applicants: Proven experience in Web Scraping (Python, BeautifulSoup, Scrapy, Apify, or similar tools). Strong background in WooCommerce data management, CSV mapping, and bulk product imports. Please provide examples of similar data extraction or e-commerce import projects you have completed. Note: We prefer to start with a milestone test of 20-30 products to ensure accuracy before proceeding with the full catalog.
I am looking exclusively for a local UAE-based developer; the total budget for this project is $10–$15, and it is non-neg...the plug-in quietly launches a web-scraper (Python) to grab supplier specs and fills the gaps before updating Airtable. Acceptance criteria 1. Signed-installer (.msi or .exe) that adds a ribbon tab in Inventor 2024+ with the above workflow. 2. Python source, neatly documented, including the computer-vision model (TensorFlow or PyTorch—your choice) and the web-scraping routines (BeautifulSoup/Scrapy). 3. REST hooks or direct connectors for Airtable and Power BI proven to work with my API keys. 4. Short video demo plus a written deployment guide. If something in the toolchain needs tweaking, I’m open to suggestions—as long as Pyt...
I need a reliable scraper that automatically captures price, reviews, and the full product description from both Amazon and Flipkart, then serves that data through a simple API I can call each day. The flow I... description) • Setup guide plus one test run proving the data refresh Acceptance criteria 1. Hitting the API after the daily run returns up-to-date fields for at least ten test products. 2. No hard-coded credentials or paths; environment variables fully supported. 3. Documentation covers installation, scaling the crawler count, and updating the SKU list. Any stack is welcome—Python (Scrapy, BeautifulSoup, Selenium), Node (Puppeteer, Cheerio), or another proven toolset—so long as it meets the requirements above and can be deployed on a standard V...
I need a reliable scraper built (Python, Scrapy / BeautifulSoup or similar) to pull data for roughly 1,000 UK companies directly from the publicly available Companies House register. For every company I want four fields captured: 1. Company name 2. Registered address 3. Full director name(s) 4. (Optional extra columns for multiple directors per company are fine) Please deliver the finished dataset as a clean, UTF-8 encoded CSV file with consistent column headers. A quick sample of 10 records up front will let me confirm the structure before you run the full job. Acceptance criteria • 1,000 distinct companies returned (no duplicates) • All mandatory fields populated; blanks only where the source truly has no data • CSV opens without formatting errors in ...
...manual trigger should be required after the first configuration. If it discovers a posting that was not present on the previous run, it should save the new record (CSV or SQLite is fine) and raise a simple desktop notification or send an email—whichever is easier for you to wire up quickly. I am comfortable with common web-automation stacks such as Python + Selenium/Playwright, BeautifulSoup, Scrapy, or even a lightweight C# solution, so feel free to choose the toolset you can ship fastest. Just package everything into a single installer or portable EXE; I do not want my end users touching the command line. Acceptance criteria: • Daily scheduled crawl without Windows Task Scheduler hacks — the app should handle timing internally. • Accurate capture...
...simple dashboard where I can start, stop, or schedule harvest runs and view basic statistics (total ads scanned, success rate, duplicate count, etc.). • Rotate user-agents, honor polite delays, and support proxy lists to minimise blocking. Captcha prompts, if encountered, should be handled automatically or flagged for manual solve. Technical notes I’m comfortable with any language—Python (Scrapy, Selenium, Playwright), C#, or Node.JS are all fine—as long as the final build runs smoothly on a standard desktop without complex setup. A compiled installer or Docker image would be ideal. Source code ownership and clear documentation are required so I can maintain or extend the tool later. Acceptance criteria • 95 % or better capture rate across Fo...
I have a running list of restaurants ...offers, and customer-facing ratings. In short, if the information is displayed on Talabat, I want it in the file. A few guardrails: • Accuracy is critical; I will spot-check random entries against the live site. • No duplicates—each restaurant should appear exactly once. • Please return the data in CSV or XLSX along with the original scrape script (Python-based preferred, using BeautifulSoup, Scrapy, or Selenium—whatever you find most reliable against Talabat’s layout). • Script must be reusable so I can refresh the dataset later without starting from scratch. If you have prior experience scraping dynamic food-delivery platforms and can deliver clean, well-documented code plus the compiled datase...
...WebSockets. Module 6: Workforce Management — Desktop monitoring agent for screenshot capture, time logging, and app/website activity tracking. Required Team Roles & Tech Stack We are hiring for the following 5 distinct roles: 1x Technical Product Manager / Project Lead: Agile sprint management, technical QA testing, code review, and milestone delivery. 1x Senior Data Extraction Engineer: Python, Scrapy, Playwright/Selenium, Proxy Rotation, Cloudflare Bypass. 1x Lead Backend Architect: Python, Django REST Framework, PostgreSQL, Django Channels (WebSockets), Redis, Celery. 1x Cross-Platform Mobile & Frontend Engineer: React.js (Web Dashboard), Flutter (iOS/Android/Desktop Chat), (Desktop Agent). 1x UI/UX Designer: Figma wireframes, component design systems, and resp...
...scraped numbers and keeps that leaderboard up to date. Please scrape all three data groups—player statistics, game results, and player profiles—then funnel them into a database. I’m comfortable using whatever engine makes the most sense (MySQL, PostgreSQL, or even SQLite); I simply need you to explain the trade-offs and set it up for me. Deliverables • A robust scraper (Python + BeautifulSoup/Scrapy/Selenium—use what’s best for the target site) • A well-structured relational database with the imported data • Clear documentation on the schema and how to refresh the data • A ranking algorithm/script that calculates player standings from the live database • A short write-up comparing the database options you consider...
I need data scraped from the Google and relevant sources. The final result should be a clean...Budget is 600 INR. Work Description: Need data of all the doctors in India. Need data of all the hospitals, clinics, and medical facilities in India. Key points I care about: • Accuracy: only valid, fully-formatted phone numbers. • Efficiency: the scrape should run unattended and respect reasonable rate limits so IPs don’t get blocked. • Reusability: deliver the source code (Python, Scrapy/BeautifulSoup/Selenium—use what you prefer) plus a brief README so I can rerun the extraction later or point it at new URLs. Let me know how you plan to tackle CAPTCHAs or dynamically loaded pages, and provide an estimated turnaround time along with a short sample ...
...capture in every run are: • product images with sponsored non sponsored split and assortment depth • inventory numbers down to each dark-store location The scraper should run headlessly, respect reasonable retry logic, evade basic bot protection with rotating fingerprints or proxy pools, and export tidy JSON or CSV files I can load straight into my dashboard. CLI-based Python is my preference (Scrapy, Playwright or a lightweight Selenium setup are all fine), but I am open to any stack that reliably delivers the data and is simple to maintain on an AWS or GCP instance. A short note on acceptance: I’ll consider the job complete once I can schedule a scrape on my own server, see all requested fields populated for at least 500 SKUs, and rerun the job 24 hours lat...
...a target or percentage-off value, and the system checks those pages frequently enough to feel “real time” without triggering the sites’ anti-bot measures. The moment a drop is detected, an email—my preferred notification channel—should land in my inbox with the new price, the old price and a direct link back to the product. I’m flexible on the tech stack, but Python (BeautifulSoup, Selenium, Scrapy), JavaScript (Puppeteer) or any other proven web-scraping framework that can handle dynamic content, rotating proxies and captchas is fine with me. A small database or flat-file store that remembers each product’s history will be useful for trend graphs later, though that’s a nice-to-have rather than a blocker. Deliverables •...
...- Optimize performance and data accuracy - Provide documentation and setup instructions Required Data Fields (examples): - Full name - Job title - Company name - Location - Industry - Profile URL - Company information - Publicly available contact information (if available) - Other custom fields based on requirements Technical Requirements: - Strong experience with: Python scraping frameworks (Scrapy, BeautifulSoup) Selenium / Playwright automation Browser automation Data processing pipelines APIs and integrations Proxy/session management Database storage Preferred Experience: - Previous LinkedIn scraping projects - Experience building lead generation tools - Experience handling large datasets - Knowledge of scraping reliability and maintenance Deliverables: - Working scraper/...
I need the full, clean dataset of every Government-run school in...Name (विद्यालय का नाम) • School Category (Primary / Upper Primary / Secondary / Higher Secondary) • Management Type (Government) No extra fields are required, and you are free to name the CSV however you like. The crucial acceptance criteria are 100 % accurate UDISE codes, zero duplicates, and full district/block coverage. Feel free to choose whichever scraping stack—BeautifulSoup, Selenium, Scrapy, or your own blend—gets the job done quickly and reliably; I care only about correctness and the 24-hour turnaround. Payment is a fixed ₹500 released on successful verification of the dataset. If you’re confident you can meet the accuracy and time demands, I’m ready to award immediat...
...scale both matter. Scope • Find only profiles that create UGC and are speaking German / are located in Germany, Austria, Switzerland. • Capture three fields per profile: email address (mandatory), public profile name/handle, and first name. • Deliver several thousand unique entries, cleaned and de-duplicated, in CSV or Google Sheet format. Technical Expectations Any combination of Python, Scrapy, Selenium, API endpoints, rotating proxies, or similar tools is fine as long as rate limits are respected and the final data is verifiable. Acceptance Criteria – Minimum agreed-upon record count met – ≤3 % bounce rate on a random sample of collected emails – No duplicate profiles across platforms – File delivered in the agreed f...
...well-structured Python script that can automatically scrape data from a target site (details shared in chat once an NDA is accepted). The script should be able to navigate through multiple pages, handle dynamic content or AJAX if present, and output clean, structured data to CSV or JSON. Please build it with widely supported libraries such as Requests/BeautifulSoup for static pages or Selenium/Scrapy if JavaScript rendering is required; I’m open to your recommendation as long as the final code is clear and fully commented. A lightweight virtual-env setup and a short README explaining how to run the script are important so I can reproduce the results on my own machine. Deliverables: • Complete, error-free Python script • or Pipfile • README with setup a...
...Seller information (username and any public contact details) • Event details (event name, venue, city, date and any section/row/seat notes) The workflow is simple: identify ticket listings across the entire U.S. market on eBay and Facebook Marketplace, extract the data points above, and deliver them in a single CSV or JSON file each run. A lightweight Python script (BeautifulSoup, Selenium, Scrapy or a comparable solution) that I can schedule on my own server would be ideal, but I’m open to alternative stacks if they meet Facebook’s and eBay’s current anti-bot measures. Acceptance criteria 1. Script runs without manual intervention and finishes a full marketplace sweep in a reasonable time. 2. Output file contains zero duplicate rows and all require...
...comments (username, karma, join date where the API allows) •any details relative to any contact information for individuals that work in SBA EIDL service department Please return everything in a clean, well-structured CSV—one file per subreddit or a single consolidated file with clear subreddit, post, and comment identifiers. A reusable script is important to me. Python with PRAW, Pushshift, Scrapy, or any Reddit-API friendly tooling is fine as long as you document the install steps and rate-limit handling so I can rerun it later without surprises. Deliverables 1. The CSV data set(s) 2. The complete, well-commented scraping script 3. A short README explaining setup, authentication, and how to rerun the job I will consider the project complete when I ...
...season performance; I’ll supply the URLs as soon as we start. The scraper must extract for every athlete: full name, weight class, team or club, most recent match outcome, cumulative win-loss record, points scored, and ranking movement over time. I want the data normalised into a single CSV and a companion JSON feed so it can drop straight into my analytics pipeline. Python is my usual stack, so Scrapy, BeautifulSoup, or a light Selenium layer for the occasional dynamic page all work. Please build in polite rate limiting, user-agent rotation, and a quick retry strategy so the job runs cleanly without stressing the sites. Deliverables • Well-documented source code with setup instructions • One-click script or scheduled task that updates the dataset automatic...
...with clean, clearly-labelled columns. A standard tabular layout is fine; no special template is required beyond what you see in the small example I have already shared. I wrote an example and explanation and have the excel example sheet ready The final deliverable is: 1. An .xlsx file containing the full dataset, ordered by model name. 2. The script or method you used (Python + BeautifulSoup/Scrapy, Power Query, etc.) so I can rerun it if the site updates. I’ll review by doing a random spot-check against the live site; every sampled figure must match what is currently displayed. If anything is missing or mis-aligned in Excel, I’ll return it for correction. Please include a short note on your chosen scraping approach and an estimated turnaround time when yo...
I’m running several CPA campaigns and I need a fresh batch of email contacts that are guaranteed to pass Debounce (or an equivalent) with a very low bounce rate. The focus is strictly on email leads sourced from two places: reputable business directories and we...for me: • You pull the data ethically and accurately, respecting each site’s terms. • Every address is verified through Debounce (or a comparable validator) before delivery. • Final output arrives in a clean CSV or Excel file with, at minimum, email, source URL, and any publicly available name or company field you can capture. If you already have an efficient scraping workflow—Python + Scrapy/Selenium, Octoparse, Apify, etc.—and can show sample rows that hit 0% invalid on Debo...
...can feed it straight into my analysis pipeline. I’ll share the exact fields during kickoff, but the scraper must be flexible enough to handle common article elements—headline, body text, author byline, publication date, and source URL—and easy to extend if I add more outlets later. Time is critical. Delivery within 24–48 hours is preferred, so please lean on a proven stack such as Python with Scrapy/BeautifulSoup, Node with Cheerio, or any robust alternative you already master. The script should: • Rotate user agents and accept a proxy list to avoid blocks • Log failed requests for easy reruns • Be clearly commented and organized so I can update selectors myself Deliverables 1. Executable script or notebook with all dependencies not...
I already have a Python-based Scrapy project and now need dedicated spiders that can harvest data from the mobile apps Rabbitmart (Egypt), Voo (Egypt) and Oscar (Egypt). The crawlers should pull every publicly available endpoint (or reverse-engineered one) required to deliver: product details, user reviews, pricing information, the underlying product IDs, any active promotions or discounts and—where the apps expose it—current stock levels. All captured information must be written to clean, well-structured JSON files so that I can feed it straight into my existing pipeline. I’m still unsure whether the apps will demand login or token-based authentication; therefore, please build the spider with the flexibility to plug in credentials, headers or session cookies if ...
Website Content Scraping Required I need content to be extracted from a website and organized in a structured format. The task includes scraping text, images (if required), and other relevant information while maintaining accuracy and proper formatting. Requirements: - Extract...accuracy and proper formatting. Requirements: - Extract content from the specified website. - Preserve headings, paragraphs, and content structure. - Organize the extracted data in Excel, CSV, or Word (as required). - Ensure the data is clean, complete, and free from duplicates. - Deliver the project within the agreed timeline. Experience with web scraping tools (such as Python, BeautifulSoup, Scrapy, Selenium, or similar) is preferred. Please mention your approach, estimated timeline, and cost in your...
...Announcement, Specialty, Experience, Experience, Book on, summary, badges and designations, (Clinic schedule, Fee, Clinic Name, Clinic Location,) of all clinics they have. Education, Med School, Residency, Fellowship Training, Certifications, online clinic hours and availability, Online clinic fee, Affiliations You may harvest the information with the tooling of your choice—Python (BeautifulSoup, Scrapy, Selenium), R, or another reliable stack—so long as the final file imports seamlessly. I am flexible on the exact format (Excel, CSV, or database dump); let me know what works best for your workflow and I’ll confirm before we start. Accuracy is critical. I will verify that every profile on the site is represented and that each required field is populat...
We're looking for an experienced Python developer to stabilize and enhance a production web scraper that's experiencing Cloudflare blocks, broken session handling, and WebSocket instability. The goal is to implement a robust, self-healing scraping pipeline with proper Redis integration. Requirements: - Strong Python experience with web scraping frameworks (Playwright, curl-cffi, Scrapy, or similar) - Hands-on experience bypassing Cloudflare challenges using TLS fingerprint matching and stealth techniques - Redis experience including session storage, TTL management, and pub-sub patterns - Experience with job queuing libraries such as BullMQ or RQ - WebSocket client implementation including reconnection logic, heartbeat management, and binary/JSON frame parsing - Ability to...
...web-scraping specialist who can reliably pull data from Facebook and Instagram. The focus is on public, real-time content; I want clean, structured results that can be analysed right away without additional tidying on my side. To be sure your approach works, please show me a short sample first—10-20 recent records from each platform will do. I am happy with any proven stack (Python + BeautifulSoup/Scrapy, Node + Puppeteer, Selenium, API work-arounds, etc.) as long as you can demonstrate stability, speed, and respect for each platform’s rate limits. Deliverables (all items required): • One working script or tool that scrapes Facebook and Instagram as agreed • A sample dataset for my review before full engagement • Clear instructions or a brief READ...
I want to build a full-featured price comparison tool similar to the core engine behind BuyHatke. The focus is strictly on comparing prices in real time, not on coupons or browser extensions. The system must work seamlessly on the web and be delivered as native iOS and A...pipeline with error handling • Relational or NoSQL store optimised for high-volume price snapshots • Clear documentation and a brief hand-off session I’ll test by verifying that identical products pulled from at least ten merchants show accurate, up-to-date prices across all three platforms within the same refresh cycle. If you’ve tackled price aggregation before—especially with tools like Python-Scrapy, Node, Firebase, or similar—you’ll be able to move quickly. Let&rsq...
...for a detail-oriented web-scraping specialist who can pull data from virtually any public-facing website and hand it back to me neatly organised in an Excel workbook. The site type and data fields will vary from project to project—sometimes it might be product listings, other times articles, reviews, or contact details—so adaptability and solid experience with tools such as Python (BeautifulSoup, Scrapy, Selenium) or equivalent are essential. What matters most is accuracy, clean formatting, and a repeatable process I can rerun in the future. Along with the finished .xlsx file, please include either the script or clear documentation of your method so I can update the scrape if the source site changes. If you have questions about pagination, login barriers, or large d...
I'm seeking a skilled Python developer to work on a project. The specific type of project, primary function, and preferred libraries/frameworks are currently undecided. Ideal Skills and Experience: - Proficiency in Python - Experience with web applications, data analysis, or automation - Familiarity with Django/Flask, Pandas/Numpy, or Scrapy/BeautifulSoup - Strong problem-solving skills - Ability to work independently and meet deadlines Please provide relevant experience and a brief project approach in your bids.
I’m kicking off a Python-based project that will either evolve into a lightweight FastAPI service or a robust web-scraping pipeline—whichever proves the better fit once we start prototyping. Clean, asynchronous code, thoughtful error handling and respect for rate limits are non-negotiable, whether you lean on FastAPI’s dependency-injection patterns or a scraping stack such as Scrapy, Playwright, Selenium, or BeautifulSoup. Before we dive into technical details, I’d like to see concrete proof of your expertise. Please share past work: a GitHub repo, a running demo, or any code snippets that highlight your mastery of FastAPI endpoints, background tasks, pydantic models, or large-scale data extraction from dynamic sites. Real examples will help me understand yo...
...across the sites. Please deliver the data in a single Google Sheets workbook, with a dedicated tab for each source site and clear column headers. So I can plan my budget, let me know your price per site along with an estimated turnaround time once you see the list and any anti-scraping measures that might require work-arounds. • If you already have tooling or scripts in Python, BeautifulSoup, Scrapy, Selenium, or similar, feel free to mention it—speed and reliability matter more to me than the specific stack....
...from a company’s website when that adds value. Core needs • A script, API, or lightweight app that inputs a keyword or industry and returns clean profile data (name, role, company, public URL). • Smart filtering to remove duplicates and obvious non-prospects. • Export options—CSV or JSON at minimum—for easy hand-off to marketing systems. Tech is up to you; Python (BeautifulSoup, Selenium, Scrapy) or a comparable stack is welcome as long as it respects LinkedIn’s limits and complies with website terms. Please send a gedetailleerd projectvoorstel outlining: 1. Your approach to bypassing anti-scraping measures without violating TOS. 2. Key milestones from prototype to final delivery. 3. Examples of similar scraping or data-collection...
I have a stream of numerical information being pulled automatically from several news sites, and I need that raw output cleaned, coded, and d...parses each scrape, identifies the key figures buried in the articles, applies consistent codes to them, and drops everything into a structured file (CSV or JSON works for me). You’ll receive: • the current scraping script and a sample of the raw dump • a field dictionary showing how each number should be labeled or categorised I’m expecting your returned script (Python preferred—BeautifulSoup/Scrapy plus pandas is perfect, but use what you like) along with the final processed dataset and a brief read-me so I can rerun or extend the pipeline later. Accuracy of the coding and reproducibility of results will be ...
...reliable, detail-oriented, and capable of handling various Python-related tasks independently. Requirements: Strong knowledge of Python Experience with Web Scraping and Data Extraction API Integration and Automation Data Processing and Data Management Ability to troubleshoot and optimize existing scripts Good communication skills Ability to meet deadlines Preferred Skills: Selenium BeautifulSoup Scrapy Pandas Requests Flask or Django (optional) Database experience (MySQL, PostgreSQL, SQLite) What I Offer: Long-term work opportunities Multiple projects every month Clear requirements and communication Prompt payment for completed work Please include: Your Python experience Examples of previous projects Your hourly rate or fixed-price expectations Your availability I am loo...
...straightforward: crawl each page, parse the HTML for the specific elements I’ll identify (headings, paragraphs, and a couple of custom tags), normalise any odd characters, then bulk-insert the results so the database is immediately query-able. A repeatable solution matters because I’ll be running the same process weekly as the sites update. I’m comfortable if you build the scraper in Python—BeautifulSoup, Scrapy, or Selenium are all fine—or you can propose another language or library you prefer, as long as it reliably handles pagination and throttles requests to stay respectful of the hosts. Deliverables: • A well-commented script or small codebase that performs the crawl, parse, and SQL insert in one run. • A SQL file (or direct push...
I need a cle...need a clean, one-time scrape of 5,000 products from 1688.com. The final file must be a CSV that contains three reliable fields for every item: the live product link, the exact title as it appears on the site, and the current price. Because 1688 is Chinese-language and often requires dynamic loading or logged-in sessions, please use whatever stack you are most comfortable with—Python, Selenium, Scrapy, Playwright, or another proven crawler—to ensure every row is complete and no data is blocked or throttled. Deliverable • A single CSV (UTF-8) containing 5,000 rows with columns: Title, Price, URL. I will verify by spot-checking random entries against the site, so accuracy and duplicate-free results are essential. Once the file passes that check the...
My day rarely looks the same twice, so I need a flexible side-kick who can switch effortlessly between deep-dive research, fast data scraping, AI-powered task execution, and polished document preparation. One hour you might be pulling product data with Python, BeautifulSoup, or Scrapy; the next you could be steering ChatGPT or another LLM to summarise findings, draft a proposal, or even spin up instructions for a third-party service I delegate to. Core responsibilities • Research – everything from quick market snapshots to more academic or product-level analysis, depending on what the week demands. • Data scraping & cleanup – locating reliable sources, extracting the essentials, and presenting them in clear, reusable formats (CSV, Google Sheets, Airtab...