245 lines
17 KiB
Markdown
245 lines
17 KiB
Markdown
# Project notes for agents working on this repo
|
|
|
|
## What this is today
|
|
|
|
A Python PoC that scrapes and writes to an HP ProCurve/OfficeConnect 1810 (J9450A)
|
|
switch's built-in web UI at `192.168.2.10` (no auth configured on the switch).
|
|
|
|
- `config.py` — connection config from env vars (`SWITCH_HOST`, `SWITCH_HTTPS`,
|
|
`SWITCH_PASSWORD`). These three map directly onto what a future Terraform provider's
|
|
config block would need per-switch.
|
|
- `switch_client.py` — session/login (`get_session()`), generic `fetch()` (GET) and
|
|
`submit_form()` (POST, for writes).
|
|
- `extractors/xe_page.py` — generic parser for the firmware's auto-generated ("XE") pages;
|
|
every page renders one of two table shapes (scalar key/value, or tabular with headers),
|
|
values always sit inside hidden `<INPUT VALUE=...>` tags. Both reads and writes rely on
|
|
this same parser.
|
|
- `extractors/*.py` — one `get_x(session)` read function per switch data page.
|
|
- `actions/*.py` — one `get_x_state(session)` / `set_x(session, value)` pair per writable
|
|
setting (currently just `locator.py`).
|
|
- `todo.py` / `TODO.md` — tracks which menu items have a read extractor implemented.
|
|
|
|
## Long-term goal: become a Terraform provider (full Go rewrite)
|
|
|
|
The intended end state is a real Terraform provider for this switch. Researched and
|
|
decided: Terraform's plugin protocol requires Go (HashiCorp's Terraform Plugin Framework
|
|
is the current recommended SDK — see sources below), and the plan is a **full rewrite in
|
|
Go**, not a thin Go shell calling out to this Python code. This Python project stays as-is
|
|
as the reference implementation the Go port gets verified against.
|
|
|
|
**Correction to an earlier assumption**: Terraform **Core** does the diffing itself
|
|
(compares saved state vs. declared config via the resource's schema) — the provider does
|
|
NOT need its own `ensure()`/diff mechanism. It only implements:
|
|
- **Read**: fetch real current state → write into Terraform state.
|
|
- **Update**: called by Core only when Core already determined there's a difference.
|
|
|
|
This still supports the existing guideline below (keep `get_x`/`set_x` separate, `set_x`
|
|
idempotent) — just for a different reason than originally assumed: not because the
|
|
provider diffs, but as a safety net against config drift (someone changes the switch via
|
|
its web UI between `terraform plan` and `apply`).
|
|
|
|
Sources: [Terraform Plugin Framework (pkg.go.dev)](https://pkg.go.dev/github.com/hashicorp/terraform-plugin-framework),
|
|
[Resources — HashiCorp Developer](https://developer.hashicorp.com/terraform/plugin/framework/resources),
|
|
[Implement a provider with the Plugin Framework (tutorial)](https://developer.hashicorp.com/terraform/tutorials/providers-plugin-framework/providers-plugin-framework-provider).
|
|
|
|
### CLI tooling: OpenTofu, not Terraform
|
|
|
|
The user is installing **OpenTofu** (`tofu` CLI) rather than the Terraform CLI. This
|
|
doesn't change any of the Go code below — OpenTofu speaks the same plugin protocol as
|
|
Terraform, and a provider built with `terraform-plugin-framework`/`terraform-plugin-go`
|
|
(there's no separate "OpenTofu SDK") works unmodified against it. Practical differences
|
|
only: the CLI binary is `tofu`, and local dev-override config lives in `~/.tofurc`
|
|
(env var `TF_CLI_CONFIG_FILE` also works) instead of `~/.terraformrc`. "Terraform
|
|
state"/"Terraform Core" below refers to the shared plugin-protocol concepts, which apply
|
|
identically under OpenTofu.
|
|
|
|
### The plan
|
|
|
|
New sibling Go module: `~/code/experiments/terraform-provider-hpe1810/` (separate
|
|
toolchain/`go.mod`, not nested in this Python project), scaffolded from HashiCorp's
|
|
`terraform-provider-scaffolding-framework` template.
|
|
|
|
- `internal/provider/client/client.go` — port of `switch_client.py` (`net/http.Client` +
|
|
`net/http/cookiejar` for session cookies; same login POST fields, same GET-before-POST
|
|
workaround, same `submit_flag=8` fix in `SubmitForm`). **Simplification vs. Python**: the
|
|
provider process lives for the whole `terraform plan`/`apply` run, so the disk-cached
|
|
session (`.session_cookie`) mostly goes away — `Configure()` logs in once, the same
|
|
client is reused for every resource/data source in that run. The switch's single-session
|
|
limit still means don't run two `terraform` processes against it concurrently.
|
|
- `internal/provider/client/xepage.go` — port of `extractors/xe_page.py`, using `goquery`
|
|
(closest Go equivalent to BeautifulSoup). Same two table shapes (scalar
|
|
`defleft`/`defright`, tabular via `<TH>`), same duplicate-header caveat as a comment.
|
|
- `provider.go` schema maps 1:1 onto `config.py`'s three fields (`host`/`https`/`password`),
|
|
with the same env var fallback pattern already established there.
|
|
- One data source per current `extractors/*.py` (`system_description`, `network_setup`,
|
|
`port_summary`, `lldp_statistics`, `buffered_log`, `mac_table`, `trunk_status`,
|
|
`loop_protection_status`, `dual_image_status`, `clock`, `sntp`, `log_configuration`,
|
|
`backup_manager`).
|
|
- One resource per current `actions/*.py` (just `locator` today) — singleton-setting
|
|
resource with a fixed state ID. `Delete` semantics are decided **per resource**, not by
|
|
one blanket rule — the guiding question is "what's the safest state for *this* setting
|
|
to end up in", which sometimes means actively changing something, not just leaving it or
|
|
resetting to a factory default:
|
|
- **Locator**: reset to `Disable` — harmless, and leaving a physical LED blinking after
|
|
`terraform destroy` would be confusing.
|
|
- **Port Configuration**: `Delete` should disable/close the port — an unmanaged port is
|
|
a security exposure, so "no longer managed by Terraform" should fail toward *safer*
|
|
(closed), not toward "whatever it happened to be".
|
|
- **Network Setup** (IP, gateway, SNMP, management VLAN): `Delete` should be a pure
|
|
no-op that only drops it from Terraform state — there's no safe "unset" state for
|
|
the switch's own management network config; touching it on delete risks the same
|
|
lockout the risk triage below already warns about.
|
|
Decide this explicitly for each resource as it's built, using this "what does *safe*
|
|
mean for this specific setting" question — don't default to a single project-wide rule.
|
|
- Testing: unit tests for `xepage.go` against `page_snapshots/*.html` copied into
|
|
`testdata/` (already-collected golden fixtures, runnable offline); acceptance tests via
|
|
`terraform-plugin-testing`, gated behind `TF_ACC=1`, only runnable against the real
|
|
switch.
|
|
- Suggested build order: scaffold + `Configure()` → one vertical slice end-to-end
|
|
(`data.hpe1810_system_description`) → remaining data sources → `resource_locator` →
|
|
further write resources in the risk-triage order below (Safe category first).
|
|
- **Status (2026-08-25)**: all 13 data sources are ported and verified with a live
|
|
`tofu plan` against the real switch (`~/code/experiments/hpe/terraform/`) — confirmed
|
|
port 24 (the active uplink) reads `Link Up`, all other ports `Link Down`.
|
|
`resource_locator` is fully built and its whole CRUD lifecycle verified live: `Create`
|
|
(enabled=true), `Update` (enabled=false), and `Delete` (`tofu destroy -target=...` ->
|
|
reset-to-Disable) all round-tripped correctly, each confirmed visually by the LED
|
|
blinking/stopping. First write resource is done end-to-end.
|
|
`resource_green_features` (Green Mode, Mode LED Time, Phy Auto Power-Down -- all three
|
|
submitted together per write, since the firmware form has no per-field submission) is
|
|
also built and its whole CRUD lifecycle verified live: `Create` matched the switch's
|
|
actual state, `Update` (mode_led_time 1->5) and `Delete` (reset to factory defaults:
|
|
Enable/10/Enable) both confirmed directly via curl against the switch (no LED to eyeball
|
|
for this one).
|
|
`resource_port` (`hpe1810_port`, `PortConfiguration.html`) is also built and its whole
|
|
CRUD lifecycle verified live, driven by `for_each` over a `var.ports` map keyed by port
|
|
number (see `~/code/experiments/hpe/terraform/variables.tf`). This page needed dedicated
|
|
handling beyond the generic parser -- see `client/port_config.go`'s doc comment for the
|
|
full writeup, short version: (1) port selection is a `submit_flag=1` reload POST with
|
|
`v_1_1_1=<port>`, not a query param; (2) the page has three duplicate-labeled "Link
|
|
Speed" rows (`v_1_8_1`/`v_1_13_1`/`v_1_15_1`, one per SFP/PHY group) that read the same
|
|
but do NOT accept the same values on write -- confirmed live, submitting all three at
|
|
once was rejected; only `v_1_8_1` (the "No SFP"/RJ45-copper field, which covers every
|
|
port on this switch) is used, and any other Physical Type is explicitly rejected rather
|
|
than guessed at; (3) three hidden context fields (`v_1_21_1` UnitIndex, `v_1_31_1`
|
|
duplicate interface, `v_1_2_1` mstid) must be included on every write or the switch
|
|
rejects with `FILTER_MISSING` -- also confirmed live, hardcoded to their observed
|
|
constant values since this is a single, non-stacked unit.
|
|
New provider-level `admin_port` attribute (int, default **24**, env fallback
|
|
`HPE1810_ADMIN_PORT`) protects the switch's uplink: `resource_port`'s `Create`/
|
|
`Update`/`Delete` all refuse to touch that interface, verified live -- a scratch resource
|
|
targeting interface 24 was correctly rejected with a diagnostic error, and port 24's live
|
|
state (`Enable`/`Auto`/`Link Up`) was confirmed unchanged afterward via curl. Full CRUD
|
|
(`Create`, `Update` speed 100->10 Mbps, `Delete` -> Disable+Auto) verified live on port 5
|
|
only, then restored to its original state (`Enable`/`100 Mbps Full Duplex`).
|
|
`resource_system_description` (`hpe1810_system_description`, `SysDescription.html`) is
|
|
also built and verified live: only 3 of the page's 9 fields are actually writable (System
|
|
Name/Location/Contact, `v_1_2_1`/`v_1_3_1`/`v_1_4_1`) -- confirmed by the page's own
|
|
client-side validation script only defining error messages for those three field ids; the
|
|
rest (Description, Software Version, Object ID, Up Time, Current Time, Date) stay
|
|
read-only, exclusively on the data source. Note this resource's TypeName intentionally
|
|
collides with the data source's (`hpe1810_system_description`) -- fine, since Terraform
|
|
keeps `resource.*`/`data.*` in separate namespaces. `Create` applied live (System Name
|
|
`Draupnir01` -> `draupnir01`), confirmed via curl. Delete resets all three fields to `""`.
|
|
Next write resources per the risk triage: Jumbo Frames, then Ping Test.
|
|
Note: switch only used on port 24 for uplink — avoid touching that port in any future
|
|
port-configuration resource testing.
|
|
- **`save_running_config` provider flag (2026-08-25)**: new optional provider attribute
|
|
(`HPE1810_SAVE_RUNNING_CONFIG` env fallback, default `false`). If `true`, `Shutdown()`
|
|
(called once at process exit, same place as the logout below) POSTs to
|
|
`SaveAllChanges.html` ("Save Configuration" in the web UI) -- but ONLY if
|
|
`client.Dirty()` is also true, i.e. some `SubmitForm` call actually succeeded this run.
|
|
`Client.dirty` is an `atomic.Bool` set inside `SubmitForm` on any successful write.
|
|
Rationale (user's, 2026-08-25): config is already source-of-truth as IaC, saving isn't
|
|
critical, and unconditionally saving on every `tofu` invocation would wear the switch's
|
|
flash for no reason -- so it only fires on runs that actually changed something, and even
|
|
then only when explicitly opted into. Verified live: `SaveAllChanges.html` accepts the
|
|
same POST shape `SubmitForm` sends (`err_flag=0` back, confirmed via curl); an
|
|
apply with `save_running_config=true` that wrote a real change completed with no errors.
|
|
- **Session logout on shutdown**: found `GET /index.html?logout=1` invalidates the session
|
|
(undocumented, discovered by probing — no `/hp_logout.html` or similar exists).
|
|
`client.Logout()` calls it; `main.go` calls `p.Logout()` right after
|
|
`providerserver.Serve` returns (process shutdown), freeing the switch's one session slot
|
|
so the *next* `tofu` invocation doesn't have to wait out the ~5 min session timeout.
|
|
Verified live: two `tofu plan` runs back-to-back now both succeed.
|
|
|
|
### What this means for code written in this Python project meanwhile
|
|
|
|
Even though the Go rewrite won't call this code directly, keep the same shape so the port
|
|
stays a straight translation:
|
|
|
|
- `get_x_state(session)` must be a pure read — no side effects, safe to call anytime.
|
|
- `set_x(session, value)` should be idempotent — calling it again with the same `value`
|
|
it already holds should be harmless. (Not yet enforced anywhere; `actions/locator.py`
|
|
always writes regardless of current state — fine for a manual on/off toggle.)
|
|
- Keep `get_x_state` and `set_x` as separate, independently callable functions — that
|
|
split is exactly what a Terraform resource's Read and Update map onto.
|
|
|
|
### Explicit non-goals right now
|
|
|
|
- No Go module has been created yet — this is still just a plan.
|
|
- No generic `ensure(desired) -> read, diff, apply-if-different` helper needed in Python
|
|
either now, given Core does the diffing in the real provider.
|
|
|
|
## Offline snapshots
|
|
|
|
`page_snapshots/` holds raw HTML + per-page `_xe_*.js` (field names/types/enums/OIDs) for
|
|
every menu page, captured while the switch was reachable, plus `current_state.json` (a
|
|
dump of every extractor's output at capture time) and `manifest.json` (HTML/JS file
|
|
mapping). Use these to write and offline-verify new `extractors/`/`actions/` modules
|
|
(parse the saved HTML directly, no live switch needed) when the switch isn't reachable —
|
|
just re-verify against the real switch once it's back online, since a snapshot can go
|
|
stale (config changes, firmware updates).
|
|
|
|
## Write-target risk triage (for future `actions/` modules)
|
|
|
|
Assessed from the snapshotted HTML/JS, not yet attempted live. Categorized by how bad a
|
|
mistake would be and how easy it is to recover from.
|
|
|
|
**Safe** (cosmetic or trivially reversible, like `actions/locator.py`):
|
|
- Green Features (`greenmode.html`) — EEE power-saving toggle, no traffic impact.
|
|
- Ping Test (`Ping.html`) — triggers a diagnostic ping, doesn't change config.
|
|
- Jumbo Frames (`JumboFrames.html`) — single global MTU toggle, reversible.
|
|
|
|
**Moderate** (affects real behavior, recoverable if you're paying attention, but a bad
|
|
value has a real consequence — treat these with a real read-before-write / dry-run, not
|
|
a blind toggle):
|
|
- LLDP Configuration, Loop Protection Config, Port Mirroring (`FDBConfig.html`),
|
|
Flow Control (`SwitchConfig.html`), Trunk Configuration/Membership.
|
|
- Time Zone / Daylight Saving Time config — wrong values just show wrong time, not
|
|
dangerous, but affects log timestamps everywhere.
|
|
|
|
**High risk — plan for a lockout, not just "recoverable"**:
|
|
- **Port Configuration** (`PortConfiguration.html`): disabling/misconfiguring the port
|
|
your management traffic actually flows over locks you out over the network (physical
|
|
access to the switch would be needed to recover).
|
|
- **VLAN Configuration / VLAN Ports / Participation** (`VlanCreateConfig.html`,
|
|
`VLANPortConfig.html`, `VLANPortParticipation.html`): misconfiguring the management
|
|
VLAN or tagging is one of the most common ways to lock yourself out of a switch
|
|
entirely.
|
|
- **Advanced Security / Secure Connection** (`security.html`, `SSLCfg.html`): e.g.
|
|
forcing HTTPS-only or an access-control change could cut off the very HTTP session
|
|
this project depends on.
|
|
|
|
**Dangerous — do not automate without explicit, per-run human confirmation, ever**:
|
|
- Reboot Switch (`SystemReset.html`) — drops all connections, however briefly.
|
|
- Factory Defaults (`ResetConfigToDefaults.html`) — wipes all configuration.
|
|
- Password Manager (`UserAccounts.html`) — setting/losing the admin password on a switch
|
|
that currently has none configured is exactly the kind of mistake that requires a
|
|
factory reset (or worse) to undo.
|
|
- Update Manager / firmware upload (`http_file_download.html`) — a bad image can brick
|
|
the switch.
|
|
- Dual Image Configuration (`dual_image_cfg.html`) — changes which firmware image boots
|
|
next; combined with a reboot this has the same blast radius as a firmware update gone
|
|
wrong.
|
|
- Save Configuration (`SaveAllChanges.html`) — not destructive by itself (persists
|
|
running-config to startup-config, standard practice on this class of switch), but it's
|
|
what turns every other write above from "revert on next reboot" into permanent —
|
|
never call this automatically as a side effect of something else.
|
|
|
|
Recommended order for future write targets, safest-value first: Green Features → Jumbo
|
|
Frames → Ping Test → LLDP/Loop Protection/Trunk config → Port Configuration → VLANs. Treat
|
|
everything in "Dangerous" as out of scope for automation entirely unless the user
|
|
explicitly asks for that specific action, and even then prefer a human-confirmed,
|
|
one-off script over a reusable `actions/` module.
|