Hello,
a while ago, there was a similar post from me, before anyone questions… however, I see myself diving deeper and deeper into the rabbit hole.
I came to this company some 4 or so months ago. It is/was in TERRIBLE state. While everything was generally working, there were frequent unexplained and unmonitored outages. The worst thing there was visibility. Why I say was…
The company used Icinga2 when I came. I did weigh in implementing a different solution, but kinda liked the idea of having a system, where I could keep the configs in the git repo and basically push with Semaphore/Ansible. But also the biggest advantage: I could let AI configure it. Besides, coming new and telling them to change the monitoring ain't really the best idea.
And that's the thing. I never ever saw Icinga2 before, nor Nagios. The company only "managed" it by adding or removing hosts, but no configuration. Configuration – and that also relatively basic stuff only, was done by external company.
Since then, I used AI HEAVILY on it. I expanded it different directions, including database monitoring, general services health, AD/DNS health checks, DHCP health, NTP, pending reboots, pending security updates, both windows and linux, event log checking, different connectivity checks etc.
While it is really great… there are descriptions to every step, there is this BIG issue: I barely understand it any more. The structure is the same, but the code in there… ufff. If someone asked why xy-sensor is not working, I would have to heavily google it or ask the external company, read and learn the documentation – for which I only have basic time for, being the only senior in the company, with a lot of pressure about other projects; or just ask the AI.
And that's my issue. One relies a LOT on AI, and I am at the point where I almost cannot troubleshoot without it! The complexity is very high, since the AI does it multiple times better than I ever will have a chance to learn.
And that doesn't only apply to monitoring. The company wants to implement IAM. They want to have it in 9 months. Many departments, and all… so they asked me whether I can automate it (or better said, I said automation is generally the key), and we fell onto Terraform. I have some knowledge of Terraform, but I know it will fall down to AI to create the scripts for EntraID etc.
There are others. Ansible-Patching (moved from AUM), config deployments, VM deployments… and Grafana/Promentheus should also come, and Kubernetes is also here (although managed by others, but partly in our hands).
I actually don't even know what question I should ask. It seems like I am way over my head. At the same time though, the job is being done and my boss isn't really keen on taking two new people to cover more. Why even? Increased stability, there are less vulnerabilities due to updates and standardization, lot of stuff just works better. So yeah, I feel like I am pulling the company from one shit (which it really really is) into another.
And now I better stop ranting, and go back to my VSCode…
Source: r/sysadmin · by /u/kosta880