使用可观测性保障安全
学习如何通过分析可观测性数据检测安全威胁和异常。了解如何为可疑活动设置告警。
使用可观测性保障安全 是 CoddyKit 上的免费 System Observability: Logging, Metrics & Tracing (ELK + OpenTelemetry) 课时。 这是第 1 节课,共 4 节。 你可以在下方免费阅读本课时的完整内容 — 然后在浏览器中使用内置代码编辑器和全天候 AI 导师进行实践。 这是 System Observability: Logging, Metrics & Tracing (ELK + OpenTelemetry) 学习路径的一部分,你的进度在网页和 CoddyKit 应用中同步。 System Observability: Logging, Metrics & Tracing (ELK + OpenTelemetry) 课程共包含 4 节课。
本课时的部分内容尚未翻译,以英文显示。
Observability for Security
Welcome! In this lesson, we'll explore how observability — our ability to understand a system from its external outputs — is a powerful tool for enhancing security.
It's not just for performance! Logs, metrics, and traces provide crucial insights into system behavior, helping us detect and respond to security threats.
Logs: Your Security Audit Trail
Logs are often the first line of defense. They record events, giving us a detailed history of what happened in a system. For security, we focus on specific types of log entries:
- Authentication: Successful and failed login attempts.
- Authorization: Changes to user permissions or access.
- Access: Attempts to access sensitive files or data.
- System Changes: Configuration updates or software installations.
- Network Events: Connection attempts, firewall blocks.
Example log entry:
{"timestamp": "2023-10-27T10:00:00Z", "event_type": "login_failed", "user": "admin", "source_ip": "192.168.1.10", "reason": "invalid_password"}Spotting Suspicious Log Patterns
By analyzing logs, we can identify patterns that often indicate malicious activity. Some common examples include:
- Brute-force attacks: Numerous failed login attempts from a single IP address or user account in a short period.
- Port scanning: Repeated connection attempts to various ports on a target system.
- Unauthorized access: Log entries showing access to resources by users without appropriate permissions.
- SQL injection attempts: Malformed database queries appearing in application logs.
Structured logging makes querying and filtering these patterns much easier!
Metrics as Security Indicators
Metrics provide aggregated data over time, which can reveal security anomalies by showing deviations from normal behavior. Look for:
- Failed Login Rate: A sudden spike could signal a brute-force attack.
- Network Traffic (Egress/Ingress): Unexpected increases might indicate data exfiltration or a Denial-of-Service (DoS) attack.
- API Error Rates: High error rates on specific endpoints, especially authorization errors (e.g., HTTP 401/403), could mean attack attempts.
- Resource Usage: Unusual spikes in CPU or memory could indicate malware, cryptominers, or unauthorized processes.
Traces for Security Context
Distributed traces track a single request as it flows through multiple services. This end-to-end view is incredibly valuable for security:
- Malicious Request Path: See the entire journey of an unauthorized request, identifying all services it touched.
- Unexpected Service Calls: Detect if a service is calling another service it shouldn't, or performing an unusual operation.
- Data Exfiltration: Trace a request that might be attempting to extract sensitive data, seeing where the data originated and where it was sent.
Traces provide the crucial context of an operation.
Alerting on Log Events
Once you know what to look for, you can set up alerts to notify you of suspicious log events. This is often done using search queries on your centralized log management system.
Examples of log-based alerts:
- Alert if
event.action: "login_failed"count exceeds 50 within 5 minutes from a singlesource.ip. - Alert if
user.role: "admin"performs anevent.action: "delete_database"outside of normal business hours. - Alert if any log contains a specific string indicating a known exploit (e.g.,
"union select password"for SQL injection).
Metric-Driven Security Alerts
Similarly, metric-based alerts can warn you when key performance indicators related to security cross certain thresholds. These alerts are great for detecting widespread or high-volume attacks.
Consider these examples:
- Alert if the
http.server.requests.status_401_totalmetric (total 401 Unauthorized responses) exceeds 100 per minute across the application. - Alert if
network.bytes_sent_totalfor the entire system increases by 200% compared to its 7-day average. - Alert if
process.cpu_usagefor an application server remains above 80% for more than 10 minutes during off-peak hours.
Correlating Signals for Deep Insights
The true power of observability for security comes from correlating all three signals: logs, metrics, and traces. No single signal tells the whole story.
- A spike in failed login metrics (metric) can trigger an investigation into specific log entries to identify the attacking IPs and usernames.
- An unusual API call observed in a trace can be cross-referenced with logs for associated errors or unauthorized attempts.
- An unauthorized access log can be linked to a trace ID to see the full path of the malicious request through your services.
This combined view enables faster and more accurate incident response.
Best Practices for Robust Security
To maximize your security posture with observability, follow these best practices:
- Granular Logging: Log enough detail to be useful, but avoid logging sensitive data directly.
- Centralized Collection: Aggregate all logs, metrics, and traces into a single, queryable platform.
- Baseline Monitoring: Understand your system's 'normal' behavior to more easily spot anomalies.
- Regular Review: Periodically audit your security alerts and dashboards to ensure they are still relevant and effective.
- Access Control: Implement least privilege for access to observability tools and data themselves.
Security Scenario Check
A user reports that their account was locked after multiple failed login attempts. Your security team suspects a brute-force attack. Which observability signals are most useful for detecting this specific type of attack and understanding its scope?
Recap: Observability for Security
Great job! You've learned how observability plays a critical role in system security. By leveraging logs, metrics, and traces, you can:
- Identify suspicious patterns and anomalies.
- Set up proactive alerts for potential threats.
- Gain deep contextual understanding during security incidents.
- Correlate data across signals for faster root cause analysis.
Integrating observability into your security strategy helps build more resilient and secure systems. Keep exploring how these powerful tools can safeguard your applications!
常见问题解答
「使用可观测性保障安全」课时是免费的吗?
是的 — 「使用可观测性保障安全」的完整文本可在网页上免费阅读。要进行交互式练习(内置代码编辑器和全天候 AI 导师)并解锁 System Observability: Logging, Metrics & Tracing (ELK + OpenTelemetry) 课程的其余内容,请升级到 CoddyKit PRO。 System Observability: Logging, Metrics & Tracing (ELK + OpenTelemetry) 课程共包含 4 节课。
「使用可观测性保障安全」这节课中我会学到什么?
学习如何通过分析可观测性数据检测安全威胁和异常。了解如何为可疑活动设置告警。 你通过在浏览器中直接运行的动手代码来练习 System Observability: Logging, Metrics & Tracing (ELK + OpenTelemetry),全天候 AI 导师会在你学习这节课的过程中回答你的问题。
学习 System Observability: Logging, Metrics & Tracing (ELK + OpenTelemetry) 需要有经验吗?
无需任何先前经验。CoddyKit 上的 System Observability: Logging, Metrics & Tracing (ELK + OpenTelemetry) 课程适合初学者到高级学习者,你可以从这里开始或从头开始,按照自己的节奏学习。 这是第 1 节课,共 4 节。
「使用可观测性保障安全」课时需要多长时间?
大多数 CoddyKit 课程大约需要 5–10 分钟。每节课都很精短且互动,所以你能稳步进步,并在网页和应用中从离开的地方继续。
我能在这节 System Observability: Logging, Metrics & Tracing (ELK + OpenTelemetry) 课中编写并运行代码吗?
能。每节 System Observability: Logging, Metrics & Tracing (ELK + OpenTelemetry) 课都包含内置代码编辑器,你可以在浏览器中直接编写并运行真实代码,并获得即时 AI 反馈 — 无需本地设置。
此课程中的所有课时
- 使用可观测性保障安全
- 性能监控与调优
- 可观测性成本优化
- 审计日志记录与合规性