Performance Engineer
Project Role Description : Diagnose issues that an in-house performance testing team has been unable to. There are five aspects to Performance Engineering: software development lifecycle and architecture, performance testing and validation, capacity planning, application performance management and problem detection and resolution.
Must have skills : Data Engineering
Good to have skills : Python (Programming Language), GitHub
Minimum 3 year(s) of experience is required
Educational Qualification : 15 years full time education
Summary:
As a Performance Engineer, a typical day involves investigating and resolving complex performance issues that have not been identified by the internal performance testing team. The role encompasses a broad spectrum of activities including analyzing software development processes and system architecture, validating performance through rigorous testing, planning for capacity needs, managing application performance, and swiftly detecting and addressing problems to ensure optimal system functionality. This position requires a proactive approach to identifying bottlenecks and collaborating with various teams to enhance overall system efficiency and reliability.
KEY RESPONSIBILITIES
Set up and configure Databricks workspaces, including cluster management, access controls, and integration with cloud identity and access management services
Design and implement Medallion architecture (Bronze, Silver, and Gold layers) on cloud object storage using Delta Lake format, ensuring data quality and traceability at each layer
Build and maintain ETL and ELT data pipelines using PySpark and Spark SQL within Databricks, covering ingestion, transformation, deduplication, normalisation, and standardisation
Ingest data from multiple source systems including relational databases (Oracle, MySQL, SQL Server), flat files, and SaaS platforms such as Salesforce, into the Lakehouse
Execute data migration from Oracle and other legacy systems, storing transformed outputs in the cloud storage Gold Layer using Databricks batch processing
Orchestrate and automate data workflows using pipeline orchestration tools such as Apache Airflow, AWS Glue, or Azure Data Factory, ensuring timely, dependency-aware, and reliable data delivery
Create and manage Linked Services, Datasets, and connection configurations for source and sink systems
Monitor pipeline execution end-to-end, troubleshoot failures, implement corrective actions, and perform root cause analysis for recurring issues
Optimise Spark jobs for performance, scalability, and cost efficiency — including partitioning strategy, caching, and resource configuration
Ensure all data is stored in Delta format with appropriate schema enforcement, ACID compliance, and incremental load patterns
Manage sensitive credentials and connection strings securely using a secrets management service (e.g., AWS Secrets Manager, Azure Key Vault, or HashiCorp Vault)
Maintain data quality standards by implementing validation and reconciliation logic within pipeline workflows
Collaborate with data analysts, data scientists, and product teams across global delivery environments to understand data needs and deliver trusted datasets
Maintain thorough documentation of pipeline designs, data flows, architecture decisions, and operational runbooks
REQUIRED SKILLS & EXPERIENCE
3–4 years of professional experience in data engineering, with a significant and demonstrable portion of that time working on Databricks in a production environment
Strong hands-on proficiency in PySpark and Python, with the ability to write clean, efficient, and maintainable transformation code
Proven experience designing and implementing Medallion architecture (Bronze / Silver / Gold) on cloud object storage, with Delta Lake as the storage layer
Solid understanding of Delta Lake capabilities: ACID transactions, schema evolution, time travel, and incremental ingestion patterns
Hands-on experience migrating and extracting data from Oracle and other relational databases (MySQL, SQL Server) into a cloud-based Lakehouse
Experience with batch data ingestion from SaaS platforms such as Salesforce
Working knowledge of at least one pipeline orchestration tool — Apache Airflow, AWS Glue, Azure Data Factory, or similar — for scheduling and dependency management
Familiarity with cloud-hosted relational databases and MySQL for structured data storage and maintenance
Sound understanding of ETL and ELT design patterns and the trade-offs between them
Proficiency in Git and GitHub for version control, branching, and collaborative code review
Experience working within Agile or Waterfall delivery frameworks with structured sprint or milestone-based planning
Strong analytical and troubleshooting skills ability to diagnose and resolve pipeline failures methodically
Effective communication and collaboration skills with the ability to work across cross-functional and geographically distributed teams
GOOD TO HAVE
Experience with Databricks Unity Catalog for data governance, lineage tracking, and fine-grained access control
Exposure to cloud secrets management integration (e.g., AWS Secrets Manager, Azure Key Vault, or HashiCorp Vault) within Databricks secret scopes
Background in legacy big data tooling — Hadoop, Hive, or Sqoop — particularly in the context of migration or modernisation projects
Familiarity with advanced Spark optimisation techniques including broadcast joins, adaptive query execution, and partition tuning
Basic understanding of data modelling concepts and dimensional or star schema design
Exposure to DevOps practices for data pipelines, including CI/CD, automated testing frameworks, or infrastructure-as-code
Databricks Certified Associate Developer for Apache Spark, or a relevant cloud data certification (AWS Certified Data Analytics, Microsoft Certified: Azure Data Engineer Associate, or equivalent)
EDUCATION
Bachelor s degree (B.Tech / B.E. or equivalent) in Computer Science, Information Technology, Data Science, or a related technical discipline is strongly preferred
A Master s degree in Computer Science, Data Engineering, or a related field is an advantage but not mandatory
Candidates with degrees in other engineering disciplines who can demonstrate strong, production-grade data engineering experience will also be considered
Databricks Certified Associate Developer for Apache Spark, or a relevant cloud data certification (AWS Certified Data Analytics, Microsoft Certified: Azure Data Engineer Associate, or equivalent) are a distinct advantage
Additional Information:
- The candidate should have minimum 3 years of experience in Data Engineering.
- This position is based at our Mumbai office.
- A 15 years full time education is required.
Navi Mumbai
Equal Employment Opportunity Statement for Australia and New Zealand
At Accenture, we recognise that our people are multi-dimensional, and we create a work environment where all people feel like they can bring their authentic selves to work, every day.
Our unwavering commitment to inclusion and diversity unleashes innovation and creates a culture where everyone feels they have equal opportunity. Our range of progressive policies support flexibility in ‘where’, ‘when’ and ‘how’ our people work to ensure that Accenture is an organisation where you can strive for more, achieve great things and maintain the balance and wellbeing you need.
We encourage applications from all people, and we are committed to removing barriers to the recruitment process and employee lifecycle. All employment decisions shall be made without regard to age, disability status, ethnicity, gender, gender identity or expression, religion or sexual orientation and we do not tolerate discrimination. If you require any accommodations or adjustments for interviews and/or at work, please reach out to exectalent@accenture.com or contact us at +61 2 9005 5000 (Australia) or +64 44666056 (New Zealand).
To ensure our workplace is inclusive and diverse we are setting bold goals and taking comprehensive action. To achieve these goals, we collect information that allows us to track the effectiveness of our Inclusion and Diversity programs. Learn how Accenture protects your personal data and know your rights in relation to your personal data. Read more about our Privacy Statement.
We work with one shared purpose: to deliver on the promise of technology and human ingenuity. Every day, more than 775,000 of us help our stakeholders continuously reinvent. Together, we drive positive change and deliver value to our clients, partners, shareholders, communities, and each other.
We believe that delivering value requires innovation, and innovation thrives in an inclusive and diverse environment. We actively foster a workplace free from bias, where everyone feels a sense of belonging and is respected and empowered to do their best work.
At Accenture, we see well-being holistically, supporting our people’s physical, mental, and financial health. We also provide opportunities to keep skills relevant through certifications, learning, and diverse work experiences. We’re proud to be consistently recognized as one of the World’s Best Workplaces™.
Join Accenture to work at the heart of change. Visit us at www.accenture.com.