Files

22 KiB

Project notes for agents working on this repo

What this is today

A Python PoC that scrapes and writes to an HP ProCurve/OfficeConnect 1810 (J9450A) switch's built-in web UI at 192.168.2.10 (no auth configured on the switch).

  • config.py — connection config from env vars (SWITCH_HOST, SWITCH_HTTPS, SWITCH_PASSWORD). These three map directly onto what a future Terraform provider's config block would need per-switch.
  • switch_client.py — session/login (get_session()), generic fetch() (GET) and submit_form() (POST, for writes).
  • extractors/xe_page.py — generic parser for the firmware's auto-generated ("XE") pages; every page renders one of two table shapes (scalar key/value, or tabular with headers), values always sit inside hidden <INPUT VALUE=...> tags. Both reads and writes rely on this same parser.
  • extractors/*.py — one get_x(session) read function per switch data page.
  • actions/*.py — one get_x_state(session) / set_x(session, value) pair per writable setting (currently just locator.py).
  • todo.py / TODO.md — tracks which menu items have a read extractor implemented.

Long-term goal: become a Terraform provider (full Go rewrite)

The intended end state is a real Terraform provider for this switch. Researched and decided: Terraform's plugin protocol requires Go (HashiCorp's Terraform Plugin Framework is the current recommended SDK — see sources below), and the plan is a full rewrite in Go, not a thin Go shell calling out to this Python code. This Python project stays as-is as the reference implementation the Go port gets verified against.

Correction to an earlier assumption: Terraform Core does the diffing itself (compares saved state vs. declared config via the resource's schema) — the provider does NOT need its own ensure()/diff mechanism. It only implements:

  • Read: fetch real current state → write into Terraform state.
  • Update: called by Core only when Core already determined there's a difference.

This still supports the existing guideline below (keep get_x/set_x separate, set_x idempotent) — just for a different reason than originally assumed: not because the provider diffs, but as a safety net against config drift (someone changes the switch via its web UI between terraform plan and apply).

Sources: Terraform Plugin Framework (pkg.go.dev), Resources — HashiCorp Developer, Implement a provider with the Plugin Framework (tutorial).

CLI tooling: OpenTofu, not Terraform

The user is installing OpenTofu (tofu CLI) rather than the Terraform CLI. This doesn't change any of the Go code below — OpenTofu speaks the same plugin protocol as Terraform, and a provider built with terraform-plugin-framework/terraform-plugin-go (there's no separate "OpenTofu SDK") works unmodified against it. Practical differences only: the CLI binary is tofu, and local dev-override config lives in ~/.tofurc (env var TF_CLI_CONFIG_FILE also works) instead of ~/.terraformrc. "Terraform state"/"Terraform Core" below refers to the shared plugin-protocol concepts, which apply identically under OpenTofu.

The plan

New sibling Go module: ~/code/experiments/terraform-provider-hpe1810/ (separate toolchain/go.mod, not nested in this Python project), scaffolded from HashiCorp's terraform-provider-scaffolding-framework template.

  • internal/provider/client/client.go — port of switch_client.py (net/http.Client + net/http/cookiejar for session cookies; same login POST fields, same GET-before-POST workaround, same submit_flag=8 fix in SubmitForm). Simplification vs. Python: the provider process lives for the whole terraform plan/apply run, so the disk-cached session (.session_cookie) mostly goes away — Configure() logs in once, the same client is reused for every resource/data source in that run. The switch's single-session limit still means don't run two terraform processes against it concurrently.
  • internal/provider/client/xepage.go — port of extractors/xe_page.py, using goquery (closest Go equivalent to BeautifulSoup). Same two table shapes (scalar defleft/defright, tabular via <TH>), same duplicate-header caveat as a comment.
  • provider.go schema maps 1:1 onto config.py's three fields (host/https/password), with the same env var fallback pattern already established there.
  • One data source per current extractors/*.py (system_description, network_setup, port_summary, lldp_statistics, buffered_log, mac_table, trunk_status, loop_protection_status, dual_image_status, clock, sntp, log_configuration, backup_manager).
  • One resource per current actions/*.py (just locator today) — singleton-setting resource with a fixed state ID. Delete semantics are decided per resource, not by one blanket rule — the guiding question is "what's the safest state for this setting to end up in", which sometimes means actively changing something, not just leaving it or resetting to a factory default:
    • Locator: reset to Disable — harmless, and leaving a physical LED blinking after terraform destroy would be confusing.
    • Port Configuration: Delete should disable/close the port — an unmanaged port is a security exposure, so "no longer managed by Terraform" should fail toward safer (closed), not toward "whatever it happened to be".
    • Network Setup (IP, gateway, SNMP, management VLAN): Delete should be a pure no-op that only drops it from Terraform state — there's no safe "unset" state for the switch's own management network config; touching it on delete risks the same lockout the risk triage below already warns about. Decide this explicitly for each resource as it's built, using this "what does safe mean for this specific setting" question — don't default to a single project-wide rule.
  • Testing: unit tests for xepage.go against page_snapshots/*.html copied into testdata/ (already-collected golden fixtures, runnable offline); acceptance tests via terraform-plugin-testing, gated behind TF_ACC=1, only runnable against the real switch.
  • Suggested build order: scaffold + Configure() → one vertical slice end-to-end (data.hpe1810_system_description) → remaining data sources → resource_locator → further write resources in the risk-triage order below (Safe category first).
  • Status (2026-08-25): all 13 data sources are ported and verified with a live tofu plan against the real switch (~/code/experiments/hpe/terraform/) — confirmed port 24 (the active uplink) reads Link Up, all other ports Link Down. resource_locator is fully built and its whole CRUD lifecycle verified live: Create (enabled=true), Update (enabled=false), and Delete (tofu destroy -target=... -> reset-to-Disable) all round-tripped correctly, each confirmed visually by the LED blinking/stopping. First write resource is done end-to-end. resource_green_features (Green Mode, Mode LED Time, Phy Auto Power-Down -- all three submitted together per write, since the firmware form has no per-field submission) is also built and its whole CRUD lifecycle verified live: Create matched the switch's actual state, Update (mode_led_time 1->5) and Delete (reset to factory defaults: Enable/10/Enable) both confirmed directly via curl against the switch (no LED to eyeball for this one). resource_port (hpe1810_port, PortConfiguration.html) is also built and its whole CRUD lifecycle verified live, driven by for_each over a var.ports map keyed by port number (see ~/code/experiments/hpe/terraform/variables.tf). This page needed dedicated handling beyond the generic parser -- see client/port_config.go's doc comment for the full writeup, short version: (1) port selection is a submit_flag=1 reload POST with v_1_1_1=<port>, not a query param; (2) the page has three duplicate-labeled "Link Speed" rows (v_1_8_1/v_1_13_1/v_1_15_1, one per SFP/PHY group) that read the same but do NOT accept the same values on write -- confirmed live, submitting all three at once was rejected; only v_1_8_1 (the "No SFP"/RJ45-copper field, which covers every port on this switch) is used, and any other Physical Type is explicitly rejected rather than guessed at; (3) three hidden context fields (v_1_21_1 UnitIndex, v_1_31_1 duplicate interface, v_1_2_1 mstid) must be included on every write or the switch rejects with FILTER_MISSING -- also confirmed live, hardcoded to their observed constant values since this is a single, non-stacked unit. New provider-level admin_port attribute (int, default 24, env fallback HPE1810_ADMIN_PORT) protects the switch's uplink: resource_port's Create/ Update/Delete all refuse to touch that interface, verified live -- a scratch resource targeting interface 24 was correctly rejected with a diagnostic error, and port 24's live state (Enable/Auto/Link Up) was confirmed unchanged afterward via curl. Full CRUD (Create, Update speed 100->10 Mbps, Delete -> Disable+Auto) verified live on port 5 only, then restored to its original state (Enable/100 Mbps Full Duplex). resource_system_description (hpe1810_system_description, SysDescription.html) is also built and verified live: only 3 of the page's 9 fields are actually writable (System Name/Location/Contact, v_1_2_1/v_1_3_1/v_1_4_1) -- confirmed by the page's own client-side validation script only defining error messages for those three field ids; the rest (Description, Software Version, Object ID, Up Time, Current Time, Date) stay read-only, exclusively on the data source. Note this resource's TypeName intentionally collides with the data source's (hpe1810_system_description) -- fine, since Terraform keeps resource.*/data.* in separate namespaces. Create applied live (System Name Draupnir01 -> draupnir01), confirmed via curl. Delete resets all three fields to "". Correction (2026-08-25): "Jumbo Frames, then Ping Test" as the suggested next targets was wrong -- neither is marked "y" in TODO.md, so neither is actually opted into. Check TODO.md's "Want it?" column before picking the next write target, not just the risk triage below (the triage says what's safe, TODO.md says what's wanted). Marked "y" but not yet implemented, as of 2026-08-25: Time Zone, Daylight Saving Time, Port Mirroring, Flow Control, Loop Protection Cfg (LoopProtectionCfg.html -- distinct from the already-done read-only Loop Protection Status), Advanced Security, Secure Connection, Trunk Configuration, Trunk Membership, LLDP Configuration, LLDP Local/Remote Device. Note: switch only used on port 24 for uplink — avoid touching that port in any future port-configuration resource testing. Update: Time Zone, Daylight Saving Time, LLDP Configuration, Flow Control, and Port Mirroring were built next per explicit user request (2026-08-25) -- see below. Remaining marked-"y"-but-not-built: Loop Protection Cfg, Advanced Security, Secure Connection, Trunk Configuration, Trunk Membership, LLDP Local/Remote Device.
  • Three more write resources (2026-08-25), user-requested: hpe1810_time_zone (Time_Zone_Configuration.html, simple 2-field scalar page, no collisions), hpe1810_lldp_configuration (LLDPConfig.html's "Global Mode" table only -- the page's per-port "Interface Mode" table, with per-port field names like 1.0.24.v_2_1_2, is a separate not-yet-built write target), and hpe1810_daylight_saving_time (Summer_Time_Configuration.html). DST was the trickiest: the page has THREE modes (Disable/Recurring/NonRecurring) each with a full field set all present in the DOM at once, several labels colliding WITHIN one table (Month/Hours/Minutes/Offset each appear in both the NonRecurring and Recurring blocks) -- handled with a dedicated regex-based field reader/writer (client/dst_config.go) rather than the generic parser, bypassing the ambiguity entirely. Confirmed live that (unlike PortConfiguration.html) this page does NOT need hidden context fields or the unused mode's fields sent -- submitting just v_1_1_1 + the Recurring block was accepted with no error. Only Disable/Recurring are supported; NonRecurring (absolute Month/Date/Year fields) is explicitly rejected. All three verified with real writes confirmed via curl (Time Zone acronym, LLDP transmit_interval, DST start_hour), then reverted to the switch's actual original values -- tofu plan shows zero drift.
  • Two more write resources (2026-08-25), user-requested, in this order: Flow Control first (hpe1810_flow_control, SwitchConfig.html, trivial single-boolean page, same shape as Locator), then Port Mirroring evaluated before building -- FDBConfig.html is misleadingly named (it's Port Mirroring, not FDB) and has two parts: a global Enable+Destination Port scalar table, and a per-source-port Direction table (25 rows: ports 1-24 + CPU, per-port field names like 1.0.25.v_1_3_2) structurally identical in complexity to LLDPConfig.html's deferred per-port table. User confirmed building only the global part as hpe1810_port_mirroring, same scope-split precedent as LLDP Configuration. Hit the same FILTER_MISSING issue as PortConfiguration.html: a hidden context field (v_2_2_2, always "1", sits next to Destination Port with no visible label) must be included on every write or the switch rejects with "Error! Failed to Set 'Enable Mirroring' ... FILTER_MISSING" -- confirmed live, fixed, then both resources verified with zero drift on tofu plan.
  • hpe1810_trunk (link aggregation / LAG), user-requested (2026-08-25): the first resource with a genuinely different shape from everything above -- a named sub-entity with real Create/Delete (TrunkConfig.html), not a singleton or a fixed for_each set. Membership assignment (TrunkMembership.html) uses a dedicated JS file (lagViewAutoGen.js) to dynamically render the picker; the raw HTML table is just a hidden data template the JS reads, not the actual submitted form -- had to read that JS to find the real field-naming convention (1.<zero-based-port-index>.24.v_2_3_1/2/3). Three write-protocol gotchas confirmed live before writing any Go code (test trunk on ports 5+6 only, never 24, cleaned up after): (1) v_2_3_1 (port number) is marked DISABLED in the raw HTML, but the page's own JS un-disables it right before submit -- a raw POST omitting it fails with FILTER_MISSING, same for the page's other hidden context fields; (2) Delete requires resubmitting that trunk's row fields (v_1_5_1..v_1_5_6) unchanged alongside the delete trigger (v_1_5_7=Enable+v_1_5_8=Delete) or it fails with "Error! Failed to remove LAG."; (3) including the page's top-level Create-section fields (v_1_1_1/v_1_2_1/v_1_4_1) in the SAME request as a delete causes the switch to create ANOTHER trunk as a side effect instead of deleting -- caught live (deleting "Trunk1" while resending v_1_4_1=Add produced a fresh empty "Trunk2"); fixed by omitting those fields entirely from delete requests. admin_mode/static_capability are exposed read-only only (Computed, not settable) -- only Create/set-members/Delete were verified live. hpe1810_trunk's members reuses the admin_port guard pattern from resource_port.go. Full CRUD verified live through the actual Terraform provider (not just curl): Create, Update (added a member), the admin_port guard (blocked live, port 24 confirmed untouched via curl afterward), and Delete -- all with zero drift on tofu plan. Multi-trunk row-indexing (1.<index>.1. prefix on TrunkConfig.html) is inferred by the same positional pattern used everywhere else in this firmware, not independently verified against a live switch with 2+ trunks (this homelab only ever had one during testing).
  • save_running_config provider flag (2026-08-25): new optional provider attribute (HPE1810_SAVE_RUNNING_CONFIG env fallback, default false). If true, Shutdown() (called once at process exit, same place as the logout below) POSTs to SaveAllChanges.html ("Save Configuration" in the web UI) -- but ONLY if client.Dirty() is also true, i.e. some SubmitForm call actually succeeded this run. Client.dirty is an atomic.Bool set inside SubmitForm on any successful write. Rationale (user's, 2026-08-25): config is already source-of-truth as IaC, saving isn't critical, and unconditionally saving on every tofu invocation would wear the switch's flash for no reason -- so it only fires on runs that actually changed something, and even then only when explicitly opted into. Verified live: SaveAllChanges.html accepts the same POST shape SubmitForm sends (err_flag=0 back, confirmed via curl); an apply with save_running_config=true that wrote a real change completed with no errors.
  • Session logout on shutdown: found GET /index.html?logout=1 invalidates the session (undocumented, discovered by probing — no /hp_logout.html or similar exists). client.Logout() calls it; main.go calls p.Logout() right after providerserver.Serve returns (process shutdown), freeing the switch's one session slot so the next tofu invocation doesn't have to wait out the ~5 min session timeout. Verified live: two tofu plan runs back-to-back now both succeed.

What this means for code written in this Python project meanwhile

Even though the Go rewrite won't call this code directly, keep the same shape so the port stays a straight translation:

  • get_x_state(session) must be a pure read — no side effects, safe to call anytime.
  • set_x(session, value) should be idempotent — calling it again with the same value it already holds should be harmless. (Not yet enforced anywhere; actions/locator.py always writes regardless of current state — fine for a manual on/off toggle.)
  • Keep get_x_state and set_x as separate, independently callable functions — that split is exactly what a Terraform resource's Read and Update map onto.

Explicit non-goals right now

  • No Go module has been created yet — this is still just a plan.
  • No generic ensure(desired) -> read, diff, apply-if-different helper needed in Python either now, given Core does the diffing in the real provider.

Offline snapshots

page_snapshots/ holds raw HTML + per-page _xe_*.js (field names/types/enums/OIDs) for every menu page, captured while the switch was reachable, plus current_state.json (a dump of every extractor's output at capture time) and manifest.json (HTML/JS file mapping). Use these to write and offline-verify new extractors//actions/ modules (parse the saved HTML directly, no live switch needed) when the switch isn't reachable — just re-verify against the real switch once it's back online, since a snapshot can go stale (config changes, firmware updates).

Write-target risk triage (for future actions/ modules)

Assessed from the snapshotted HTML/JS, not yet attempted live. Categorized by how bad a mistake would be and how easy it is to recover from.

Safe (cosmetic or trivially reversible, like actions/locator.py):

  • Green Features (greenmode.html) — EEE power-saving toggle, no traffic impact.
  • Ping Test (Ping.html) — triggers a diagnostic ping, doesn't change config.
  • Jumbo Frames (JumboFrames.html) — single global MTU toggle, reversible.

Moderate (affects real behavior, recoverable if you're paying attention, but a bad value has a real consequence — treat these with a real read-before-write / dry-run, not a blind toggle):

  • LLDP Configuration, Loop Protection Config, Port Mirroring (FDBConfig.html), Flow Control (SwitchConfig.html), Trunk Configuration/Membership.
  • Time Zone / Daylight Saving Time config — wrong values just show wrong time, not dangerous, but affects log timestamps everywhere.

High risk — plan for a lockout, not just "recoverable":

  • Port Configuration (PortConfiguration.html): disabling/misconfiguring the port your management traffic actually flows over locks you out over the network (physical access to the switch would be needed to recover).
  • VLAN Configuration / VLAN Ports / Participation (VlanCreateConfig.html, VLANPortConfig.html, VLANPortParticipation.html): misconfiguring the management VLAN or tagging is one of the most common ways to lock yourself out of a switch entirely.
  • Advanced Security / Secure Connection (security.html, SSLCfg.html): e.g. forcing HTTPS-only or an access-control change could cut off the very HTTP session this project depends on.

Dangerous — do not automate without explicit, per-run human confirmation, ever:

  • Reboot Switch (SystemReset.html) — drops all connections, however briefly.
  • Factory Defaults (ResetConfigToDefaults.html) — wipes all configuration.
  • Password Manager (UserAccounts.html) — setting/losing the admin password on a switch that currently has none configured is exactly the kind of mistake that requires a factory reset (or worse) to undo.
  • Update Manager / firmware upload (http_file_download.html) — a bad image can brick the switch.
  • Dual Image Configuration (dual_image_cfg.html) — changes which firmware image boots next; combined with a reboot this has the same blast radius as a firmware update gone wrong.
  • Save Configuration (SaveAllChanges.html) — not destructive by itself (persists running-config to startup-config, standard practice on this class of switch), but it's what turns every other write above from "revert on next reboot" into permanent — never call this automatically as a side effect of something else.

Recommended order for future write targets, safest-value first: Green Features → Jumbo Frames → Ping Test → LLDP/Loop Protection/Trunk config → Port Configuration → VLANs. Treat everything in "Dangerous" as out of scope for automation entirely unless the user explicitly asks for that specific action, and even then prefer a human-confirmed, one-off script over a reusable actions/ module.