Files
hp-iac/AGENT.md
T

314 lines
22 KiB
Markdown

# Project notes for agents working on this repo
## What this is today
A Python PoC that scrapes and writes to an HP ProCurve/OfficeConnect 1810 (J9450A)
switch's built-in web UI at `192.168.2.10` (no auth configured on the switch).
- `config.py` — connection config from env vars (`SWITCH_HOST`, `SWITCH_HTTPS`,
`SWITCH_PASSWORD`). These three map directly onto what a future Terraform provider's
config block would need per-switch.
- `switch_client.py` — session/login (`get_session()`), generic `fetch()` (GET) and
`submit_form()` (POST, for writes).
- `extractors/xe_page.py` — generic parser for the firmware's auto-generated ("XE") pages;
every page renders one of two table shapes (scalar key/value, or tabular with headers),
values always sit inside hidden `<INPUT VALUE=...>` tags. Both reads and writes rely on
this same parser.
- `extractors/*.py` — one `get_x(session)` read function per switch data page.
- `actions/*.py` — one `get_x_state(session)` / `set_x(session, value)` pair per writable
setting (currently just `locator.py`).
- `todo.py` / `TODO.md` — tracks which menu items have a read extractor implemented.
## Long-term goal: become a Terraform provider (full Go rewrite)
The intended end state is a real Terraform provider for this switch. Researched and
decided: Terraform's plugin protocol requires Go (HashiCorp's Terraform Plugin Framework
is the current recommended SDK — see sources below), and the plan is a **full rewrite in
Go**, not a thin Go shell calling out to this Python code. This Python project stays as-is
as the reference implementation the Go port gets verified against.
**Correction to an earlier assumption**: Terraform **Core** does the diffing itself
(compares saved state vs. declared config via the resource's schema) — the provider does
NOT need its own `ensure()`/diff mechanism. It only implements:
- **Read**: fetch real current state → write into Terraform state.
- **Update**: called by Core only when Core already determined there's a difference.
This still supports the existing guideline below (keep `get_x`/`set_x` separate, `set_x`
idempotent) — just for a different reason than originally assumed: not because the
provider diffs, but as a safety net against config drift (someone changes the switch via
its web UI between `terraform plan` and `apply`).
Sources: [Terraform Plugin Framework (pkg.go.dev)](https://pkg.go.dev/github.com/hashicorp/terraform-plugin-framework),
[Resources — HashiCorp Developer](https://developer.hashicorp.com/terraform/plugin/framework/resources),
[Implement a provider with the Plugin Framework (tutorial)](https://developer.hashicorp.com/terraform/tutorials/providers-plugin-framework/providers-plugin-framework-provider).
### CLI tooling: OpenTofu, not Terraform
The user is installing **OpenTofu** (`tofu` CLI) rather than the Terraform CLI. This
doesn't change any of the Go code below — OpenTofu speaks the same plugin protocol as
Terraform, and a provider built with `terraform-plugin-framework`/`terraform-plugin-go`
(there's no separate "OpenTofu SDK") works unmodified against it. Practical differences
only: the CLI binary is `tofu`, and local dev-override config lives in `~/.tofurc`
(env var `TF_CLI_CONFIG_FILE` also works) instead of `~/.terraformrc`. "Terraform
state"/"Terraform Core" below refers to the shared plugin-protocol concepts, which apply
identically under OpenTofu.
### The plan
New sibling Go module: `~/code/experiments/terraform-provider-hpe1810/` (separate
toolchain/`go.mod`, not nested in this Python project), scaffolded from HashiCorp's
`terraform-provider-scaffolding-framework` template.
- `internal/provider/client/client.go` — port of `switch_client.py` (`net/http.Client` +
`net/http/cookiejar` for session cookies; same login POST fields, same GET-before-POST
workaround, same `submit_flag=8` fix in `SubmitForm`). **Simplification vs. Python**: the
provider process lives for the whole `terraform plan`/`apply` run, so the disk-cached
session (`.session_cookie`) mostly goes away — `Configure()` logs in once, the same
client is reused for every resource/data source in that run. The switch's single-session
limit still means don't run two `terraform` processes against it concurrently.
- `internal/provider/client/xepage.go` — port of `extractors/xe_page.py`, using `goquery`
(closest Go equivalent to BeautifulSoup). Same two table shapes (scalar
`defleft`/`defright`, tabular via `<TH>`), same duplicate-header caveat as a comment.
- `provider.go` schema maps 1:1 onto `config.py`'s three fields (`host`/`https`/`password`),
with the same env var fallback pattern already established there.
- One data source per current `extractors/*.py` (`system_description`, `network_setup`,
`port_summary`, `lldp_statistics`, `buffered_log`, `mac_table`, `trunk_status`,
`loop_protection_status`, `dual_image_status`, `clock`, `sntp`, `log_configuration`,
`backup_manager`).
- One resource per current `actions/*.py` (just `locator` today) — singleton-setting
resource with a fixed state ID. `Delete` semantics are decided **per resource**, not by
one blanket rule — the guiding question is "what's the safest state for *this* setting
to end up in", which sometimes means actively changing something, not just leaving it or
resetting to a factory default:
- **Locator**: reset to `Disable` — harmless, and leaving a physical LED blinking after
`terraform destroy` would be confusing.
- **Port Configuration**: `Delete` should disable/close the port — an unmanaged port is
a security exposure, so "no longer managed by Terraform" should fail toward *safer*
(closed), not toward "whatever it happened to be".
- **Network Setup** (IP, gateway, SNMP, management VLAN): `Delete` should be a pure
no-op that only drops it from Terraform state — there's no safe "unset" state for
the switch's own management network config; touching it on delete risks the same
lockout the risk triage below already warns about.
Decide this explicitly for each resource as it's built, using this "what does *safe*
mean for this specific setting" question — don't default to a single project-wide rule.
- Testing: unit tests for `xepage.go` against `page_snapshots/*.html` copied into
`testdata/` (already-collected golden fixtures, runnable offline); acceptance tests via
`terraform-plugin-testing`, gated behind `TF_ACC=1`, only runnable against the real
switch.
- Suggested build order: scaffold + `Configure()` → one vertical slice end-to-end
(`data.hpe1810_system_description`) → remaining data sources → `resource_locator` →
further write resources in the risk-triage order below (Safe category first).
- **Status (2026-08-25)**: all 13 data sources are ported and verified with a live
`tofu plan` against the real switch (`~/code/experiments/hpe/terraform/`) — confirmed
port 24 (the active uplink) reads `Link Up`, all other ports `Link Down`.
`resource_locator` is fully built and its whole CRUD lifecycle verified live: `Create`
(enabled=true), `Update` (enabled=false), and `Delete` (`tofu destroy -target=...` ->
reset-to-Disable) all round-tripped correctly, each confirmed visually by the LED
blinking/stopping. First write resource is done end-to-end.
`resource_green_features` (Green Mode, Mode LED Time, Phy Auto Power-Down -- all three
submitted together per write, since the firmware form has no per-field submission) is
also built and its whole CRUD lifecycle verified live: `Create` matched the switch's
actual state, `Update` (mode_led_time 1->5) and `Delete` (reset to factory defaults:
Enable/10/Enable) both confirmed directly via curl against the switch (no LED to eyeball
for this one).
`resource_port` (`hpe1810_port`, `PortConfiguration.html`) is also built and its whole
CRUD lifecycle verified live, driven by `for_each` over a `var.ports` map keyed by port
number (see `~/code/experiments/hpe/terraform/variables.tf`). This page needed dedicated
handling beyond the generic parser -- see `client/port_config.go`'s doc comment for the
full writeup, short version: (1) port selection is a `submit_flag=1` reload POST with
`v_1_1_1=<port>`, not a query param; (2) the page has three duplicate-labeled "Link
Speed" rows (`v_1_8_1`/`v_1_13_1`/`v_1_15_1`, one per SFP/PHY group) that read the same
but do NOT accept the same values on write -- confirmed live, submitting all three at
once was rejected; only `v_1_8_1` (the "No SFP"/RJ45-copper field, which covers every
port on this switch) is used, and any other Physical Type is explicitly rejected rather
than guessed at; (3) three hidden context fields (`v_1_21_1` UnitIndex, `v_1_31_1`
duplicate interface, `v_1_2_1` mstid) must be included on every write or the switch
rejects with `FILTER_MISSING` -- also confirmed live, hardcoded to their observed
constant values since this is a single, non-stacked unit.
New provider-level `admin_port` attribute (int, default **24**, env fallback
`HPE1810_ADMIN_PORT`) protects the switch's uplink: `resource_port`'s `Create`/
`Update`/`Delete` all refuse to touch that interface, verified live -- a scratch resource
targeting interface 24 was correctly rejected with a diagnostic error, and port 24's live
state (`Enable`/`Auto`/`Link Up`) was confirmed unchanged afterward via curl. Full CRUD
(`Create`, `Update` speed 100->10 Mbps, `Delete` -> Disable+Auto) verified live on port 5
only, then restored to its original state (`Enable`/`100 Mbps Full Duplex`).
`resource_system_description` (`hpe1810_system_description`, `SysDescription.html`) is
also built and verified live: only 3 of the page's 9 fields are actually writable (System
Name/Location/Contact, `v_1_2_1`/`v_1_3_1`/`v_1_4_1`) -- confirmed by the page's own
client-side validation script only defining error messages for those three field ids; the
rest (Description, Software Version, Object ID, Up Time, Current Time, Date) stay
read-only, exclusively on the data source. Note this resource's TypeName intentionally
collides with the data source's (`hpe1810_system_description`) -- fine, since Terraform
keeps `resource.*`/`data.*` in separate namespaces. `Create` applied live (System Name
`Draupnir01` -> `draupnir01`), confirmed via curl. Delete resets all three fields to `""`.
Correction (2026-08-25): "Jumbo Frames, then Ping Test" as the suggested next targets
was wrong -- neither is marked "y" in TODO.md, so neither is actually opted into. Check
TODO.md's "Want it?" column before picking the next write target, not just the risk
triage below (the triage says what's *safe*, TODO.md says what's *wanted*). Marked "y"
but not yet implemented, as of 2026-08-25: Time Zone, Daylight Saving Time, Port
Mirroring, Flow Control, Loop Protection Cfg (LoopProtectionCfg.html -- distinct from the
already-done read-only Loop Protection Status), Advanced Security, Secure Connection,
Trunk Configuration, Trunk Membership, LLDP Configuration, LLDP Local/Remote Device.
Note: switch only used on port 24 for uplink — avoid touching that port in any future
port-configuration resource testing.
**Update**: Time Zone, Daylight Saving Time, LLDP Configuration, Flow Control, and Port
Mirroring were built next per explicit user request (2026-08-25) -- see below. Remaining
marked-"y"-but-not-built: Loop Protection Cfg, Advanced Security, Secure Connection,
Trunk Configuration, Trunk Membership, LLDP Local/Remote Device.
- Three more write resources (2026-08-25), user-requested: `hpe1810_time_zone`
(`Time_Zone_Configuration.html`, simple 2-field scalar page, no collisions),
`hpe1810_lldp_configuration` (`LLDPConfig.html`'s "Global Mode" table only -- the page's
per-port "Interface Mode" table, with per-port field names like `1.0.24.v_2_1_2`, is a
separate not-yet-built write target), and `hpe1810_daylight_saving_time`
(`Summer_Time_Configuration.html`). DST was the trickiest: the page has THREE modes
(Disable/Recurring/NonRecurring) each with a full field set all present in the DOM at
once, several labels colliding WITHIN one table (Month/Hours/Minutes/Offset each appear
in both the NonRecurring and Recurring blocks) -- handled with a dedicated regex-based
field reader/writer (`client/dst_config.go`) rather than the generic parser, bypassing
the ambiguity entirely. Confirmed live that (unlike PortConfiguration.html) this page
does NOT need hidden context fields or the unused mode's fields sent -- submitting just
`v_1_1_1` + the Recurring block was accepted with no error. Only Disable/Recurring are
supported; NonRecurring (absolute Month/Date/Year fields) is explicitly rejected. All
three verified with real writes confirmed via curl (Time Zone acronym, LLDP
transmit_interval, DST start_hour), then reverted to the switch's actual original values
-- `tofu plan` shows zero drift.
- Two more write resources (2026-08-25), user-requested, in this order: Flow Control
first (`hpe1810_flow_control`, `SwitchConfig.html`, trivial single-boolean page, same
shape as Locator), then Port Mirroring evaluated before building -- `FDBConfig.html` is
misleadingly named (it's Port Mirroring, not FDB) and has two parts: a global
Enable+Destination Port scalar table, and a per-source-port Direction table (25 rows:
ports 1-24 + CPU, per-port field names like `1.0.25.v_1_3_2`) structurally identical in
complexity to LLDPConfig.html's deferred per-port table. User confirmed building only the
global part as `hpe1810_port_mirroring`, same scope-split precedent as LLDP
Configuration. Hit the same `FILTER_MISSING` issue as PortConfiguration.html: a hidden
context field (`v_2_2_2`, always `"1"`, sits next to Destination Port with no visible
label) must be included on every write or the switch rejects with "Error! Failed to Set
'Enable Mirroring' ... FILTER_MISSING" -- confirmed live, fixed, then both resources
verified with zero drift on `tofu plan`.
- **`hpe1810_trunk` (link aggregation / LAG), user-requested (2026-08-25)**: the first
resource with a genuinely different shape from everything above -- a named sub-entity
with real Create/Delete (`TrunkConfig.html`), not a singleton or a fixed for_each set.
Membership assignment (`TrunkMembership.html`) uses a dedicated JS file
(`lagViewAutoGen.js`) to dynamically render the picker; the raw HTML table is just a
hidden data template the JS reads, not the actual submitted form -- had to read that JS
to find the real field-naming convention (`1.<zero-based-port-index>.24.v_2_3_1/2/3`).
Three write-protocol gotchas confirmed live before writing any Go code (test trunk on
ports 5+6 only, never 24, cleaned up after): (1) `v_2_3_1` (port number) is marked
`DISABLED` in the raw HTML, but the page's own JS un-disables it right before submit --
a raw POST omitting it fails with `FILTER_MISSING`, same for the page's other hidden
context fields; (2) Delete requires resubmitting that trunk's row fields
(`v_1_5_1`..`v_1_5_6`) unchanged alongside the delete trigger
(`v_1_5_7=Enable`+`v_1_5_8=Delete`) or it fails with "Error! Failed to remove LAG.";
(3) **including the page's top-level Create-section fields
(`v_1_1_1`/`v_1_2_1`/`v_1_4_1`) in the SAME request as a delete causes the switch to
create ANOTHER trunk as a side effect instead of deleting** -- caught live (deleting
"Trunk1" while resending `v_1_4_1=Add` produced a fresh empty "Trunk2"); fixed by
omitting those fields entirely from delete requests. `admin_mode`/`static_capability`
are exposed read-only only (Computed, not settable) -- only Create/set-members/Delete
were verified live. `hpe1810_trunk`'s `members` reuses the `admin_port` guard pattern
from `resource_port.go`. Full CRUD verified live **through the actual Terraform
provider** (not just curl): Create, Update (added a member), the admin_port guard
(blocked live, port 24 confirmed untouched via curl afterward), and Delete -- all with
zero drift on `tofu plan`. Multi-trunk row-indexing (`1.<index>.1.` prefix on
TrunkConfig.html) is inferred by the same positional pattern used everywhere else in
this firmware, not independently verified against a live switch with 2+ trunks (this
homelab only ever had one during testing).
- **`save_running_config` provider flag (2026-08-25)**: new optional provider attribute
(`HPE1810_SAVE_RUNNING_CONFIG` env fallback, default `false`). If `true`, `Shutdown()`
(called once at process exit, same place as the logout below) POSTs to
`SaveAllChanges.html` ("Save Configuration" in the web UI) -- but ONLY if
`client.Dirty()` is also true, i.e. some `SubmitForm` call actually succeeded this run.
`Client.dirty` is an `atomic.Bool` set inside `SubmitForm` on any successful write.
Rationale (user's, 2026-08-25): config is already source-of-truth as IaC, saving isn't
critical, and unconditionally saving on every `tofu` invocation would wear the switch's
flash for no reason -- so it only fires on runs that actually changed something, and even
then only when explicitly opted into. Verified live: `SaveAllChanges.html` accepts the
same POST shape `SubmitForm` sends (`err_flag=0` back, confirmed via curl); an
apply with `save_running_config=true` that wrote a real change completed with no errors.
- **Session logout on shutdown**: found `GET /index.html?logout=1` invalidates the session
(undocumented, discovered by probing — no `/hp_logout.html` or similar exists).
`client.Logout()` calls it; `main.go` calls `p.Logout()` right after
`providerserver.Serve` returns (process shutdown), freeing the switch's one session slot
so the *next* `tofu` invocation doesn't have to wait out the ~5 min session timeout.
Verified live: two `tofu plan` runs back-to-back now both succeed.
### What this means for code written in this Python project meanwhile
Even though the Go rewrite won't call this code directly, keep the same shape so the port
stays a straight translation:
- `get_x_state(session)` must be a pure read — no side effects, safe to call anytime.
- `set_x(session, value)` should be idempotent — calling it again with the same `value`
it already holds should be harmless. (Not yet enforced anywhere; `actions/locator.py`
always writes regardless of current state — fine for a manual on/off toggle.)
- Keep `get_x_state` and `set_x` as separate, independently callable functions — that
split is exactly what a Terraform resource's Read and Update map onto.
### Explicit non-goals right now
- No Go module has been created yet — this is still just a plan.
- No generic `ensure(desired) -> read, diff, apply-if-different` helper needed in Python
either now, given Core does the diffing in the real provider.
## Offline snapshots
`page_snapshots/` holds raw HTML + per-page `_xe_*.js` (field names/types/enums/OIDs) for
every menu page, captured while the switch was reachable, plus `current_state.json` (a
dump of every extractor's output at capture time) and `manifest.json` (HTML/JS file
mapping). Use these to write and offline-verify new `extractors/`/`actions/` modules
(parse the saved HTML directly, no live switch needed) when the switch isn't reachable —
just re-verify against the real switch once it's back online, since a snapshot can go
stale (config changes, firmware updates).
## Write-target risk triage (for future `actions/` modules)
Assessed from the snapshotted HTML/JS, not yet attempted live. Categorized by how bad a
mistake would be and how easy it is to recover from.
**Safe** (cosmetic or trivially reversible, like `actions/locator.py`):
- Green Features (`greenmode.html`) — EEE power-saving toggle, no traffic impact.
- Ping Test (`Ping.html`) — triggers a diagnostic ping, doesn't change config.
- Jumbo Frames (`JumboFrames.html`) — single global MTU toggle, reversible.
**Moderate** (affects real behavior, recoverable if you're paying attention, but a bad
value has a real consequence — treat these with a real read-before-write / dry-run, not
a blind toggle):
- LLDP Configuration, Loop Protection Config, Port Mirroring (`FDBConfig.html`),
Flow Control (`SwitchConfig.html`), Trunk Configuration/Membership.
- Time Zone / Daylight Saving Time config — wrong values just show wrong time, not
dangerous, but affects log timestamps everywhere.
**High risk — plan for a lockout, not just "recoverable"**:
- **Port Configuration** (`PortConfiguration.html`): disabling/misconfiguring the port
your management traffic actually flows over locks you out over the network (physical
access to the switch would be needed to recover).
- **VLAN Configuration / VLAN Ports / Participation** (`VlanCreateConfig.html`,
`VLANPortConfig.html`, `VLANPortParticipation.html`): misconfiguring the management
VLAN or tagging is one of the most common ways to lock yourself out of a switch
entirely.
- **Advanced Security / Secure Connection** (`security.html`, `SSLCfg.html`): e.g.
forcing HTTPS-only or an access-control change could cut off the very HTTP session
this project depends on.
**Dangerous — do not automate without explicit, per-run human confirmation, ever**:
- Reboot Switch (`SystemReset.html`) — drops all connections, however briefly.
- Factory Defaults (`ResetConfigToDefaults.html`) — wipes all configuration.
- Password Manager (`UserAccounts.html`) — setting/losing the admin password on a switch
that currently has none configured is exactly the kind of mistake that requires a
factory reset (or worse) to undo.
- Update Manager / firmware upload (`http_file_download.html`) — a bad image can brick
the switch.
- Dual Image Configuration (`dual_image_cfg.html`) — changes which firmware image boots
next; combined with a reboot this has the same blast radius as a firmware update gone
wrong.
- Save Configuration (`SaveAllChanges.html`) — not destructive by itself (persists
running-config to startup-config, standard practice on this class of switch), but it's
what turns every other write above from "revert on next reboot" into permanent —
never call this automatically as a side effect of something else.
Recommended order for future write targets, safest-value first: Green Features → Jumbo
Frames → Ping Test → LLDP/Loop Protection/Trunk config → Port Configuration → VLANs. Treat
everything in "Dangerous" as out of scope for automation entirely unless the user
explicitly asks for that specific action, and even then prefer a human-confirmed,
one-off script over a reusable `actions/` module.