Description
Job Summary:
We are seeking a Senior Technology Engineer specialized in Observability to lead the evolution of observability capabilities at a leading bank, strengthening incident detection and system reliability.
Key Highlights:
1. Lead the technical evolution of observability at a leading bank.
2. Design and implement observability solutions for complex platforms.
3. Serve as the technical reference point for observability and reliability.
DESCRIPTION
### **About the Opportunity:**
Our client is a **leader in the Banking Sector**, with one of the largest and most complex technological ecosystems in the country: cloud platforms, data centers, mainframes, and mission-critical applications serving millions of customers daily.
They seek to hire a **Senior Technology Engineer specialized in Observability**, whose key challenge will be to **technically lead the evolution of the bank’s observability capabilities**, ensuring integration, processing, and visualization of metrics, logs, events, and distributed traces originating from multiple technology platforms.
The objective is to enable an **end\-to\-end view of application, infrastructure, and business service health**, strengthening early incident detection, root cause analysis, and reliability of critical services.
### **How Will You Make a Big Impact?**
* Technically lead the integration of new data sources into the corporate observability ecosystem.
* Design and implement solutions for capturing, processing, and visualizing metrics, logs, events, and distributed traces.
* Administer and evolve observability platforms monitoring critical services.
* Define observability standards and best practices for various technology teams.
* Design comprehensive monitoring mechanisms for applications, infrastructure, and business services.
* Participate in complex incident analysis, identifying root causes and proposing permanent improvements.
* Collaborate with Architecture, Development, Infrastructure, Cloud, and SRE teams to strengthen monitoring and reliability capabilities.
* Serve as the technical reference point for internal teams and specialized vendors.
* Provide technical advisory support for decision-making and evolution of observability solutions.
* Drive automation initiatives, proactive monitoring, and continuous improvement.
REQUIREMENTS
**Experience:**
* Over **5 years of experience** in technology roles related to observability, monitoring, SRE, platforms, reliability, or infrastructure.
* Minimum **3 years of experience working with enterprise observability platforms**.
* Experience implementing observability and monitoring solutions in complex, distributed environments.
* Experience integrating data from multiple platforms and technologies.
* Experience resolving, diagnosing, and analyzing high-severity incidents.
* Experience exercising **technical leadership**, advising teams or squads, and participating in defining technical solutions and strategies.
**Essential Technical Knowledge:**
* **Strong and advanced experience with at least one** of the following platforms: **Grafana, Dynatrace, or Datadog**.
* Advanced knowledge of **OpenTelemetry (OTEL)**: application instrumentation, metric collection, log management, distributed tracing, and telemetry integration in distributed architectures.
* Experience in **Enterprise Observability** and **Application Performance Monitoring (APM)**.
* Metrics, logs, and distributed traces.
* Root Cause Analysis (RCA).
* APIs and integrations.
* Cloud Computing (AWS, Azure, or GCP).
* Kubernetes and containers.
* Linux.
* Distributed architectures and microservices.
**Education:**
* Bachelor’s degree or graduate in Systems Engineering, Computer Engineering, Software Engineering, Electronic Engineering, or related fields.
#### **Desirable:**
* Experience in the financial sector or high-transaction environments.
* Knowledge of **Prometheus, ELK**, or other complementary observability tools.
* Experience in automation via scripting (**Python, Bash, or others**).