More jobs:
Job Description & How to Apply Below
This is an opportunity to join a team-first meritocracy and help grow an entrepreneurial group inside Alternative Path. You will be asked to contribute, given ownership, and expected to make your voice heard.
Role Summary
Perform web scraping using various scraping techniques and then utilize Python's Pandas library for data cleaning and manipulation. Then ingest the data into a Database/Warehouse and schedule the scrapers using Airflow or other tools.
Requirements
Role Overview
The Web Scraping Team at Alternative Path is seeking a creative and detail-oriented developer to contribute to client projects. The team develops essential applications, datasets, and alerts for various teams within the client's organization, supporting their daily investment decisions. The mission is to maintain operational excellence by delivering high-quality proprietary datasets, timely notifications, and exceptional service. We are seeking someone self-motivated, self-sufficient, with a passion for tinkering and a love for automation.
In your role, you will:
- Collaborate with analysts to understand and anticipate requirements.
- Design, implement, and maintain web scrapers for a wide variety of alternative datasets.
- Perform data cleaning, Exploration, transformation, etc., of scraped data.
- Collaborate with cross-functional teams to understand data requirements and implement efficient data processing workflows.
- Author QC checks to validate data availability and integrity. Maintain alerting systems and investigate time-sensitive data incidents to ensure smooth day- today operations.
- Design and implement products and tools to enhance the Web scraping Platform.
Qualifications
Must have
- Bachelors/master's degree in computer science or in any related field
- 4-6 years of experience in the data engineering field
- Strong Python and SQL/Database skills
- Strong expertise in using the Pandas library (Python) is a must
- Strong expertise with scraping and common scraping tools(Selenium, Scrappy, Fiddler, Postman, XPath)
- Strong expertise in using the Apache Airflow platform for workflow orchestration and data pipeline management is a must.
- Experience with web technologies(HTML/JS, APIs, etc.)
- Proven work experience in working with large data sets for Data cleaning, Data transformation, Data manipulation, and Data replacements.
- Excellent verbal and written communication skills
- Aptitude for designing infrastructure, data products, and tools for Data Scientists
Preferred
- Experience containerizing workloads with Docker(Kubernetes a plus)
- Experience with building automation (Jenkins, Git Hub Actions, CI/CD)
- Experience with AWS technologies like S3, RDS, SNS, SQS, Lambda, etc
Position Requirements
10+ Years
work experience
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
Search for further Jobs Here:
×