Вход на сайт

Просмотр новости

Найдите то, что Вас интересует

Recovering a Morpheus Appliance from MySQL InnoDB Corruption

Дата публикации: 26-09-2026 13:53:22

A Morpheus appliance landed on my plate in bad shape. MySQL was crash-looping, the UI was dead, and nobody could provision anything. The embedded database had picked up InnoDB corruption, and InnoDB was refusing to start rather than hand back data it couldn't vouch for - and the worst part of it all - default Morpheus backup was disabled Note: This blog is about how I got it back to a working state, and not to be referred to as any official documentation. I've kept the reasoning behind each decision in here, along withthe things that went wrong halfway through, because those were the parts I couldn't find documented anywhere when I actually needed them. Final count: 35,687 rows out of 35,687,with exactly one corrupted field lost. WHAT WAS THE INNODB CORRUPTION?-InnoDB is MySQL's default transactional storage engine. Two internal logs keep it honest:Redo log - records changes so they can be replayed after a crash Undo log (tablespace) - records how to roll back uncommitted transactionsEvery data page carries a Log Sequence Number (LSN), which is basically a version stamp for the last change made to that page. On startup, InnoDB replays the redo log until every page has caught up to the current LSN.The trouble starts when a page claims an LSN higher than the redo log itself. That should not be possible. InnoDB can't reconcile it and can't prove the page is valid, so it does the only safe thing available and aborts. This was the part which seemed to be the main issue lying around. THE INCIDENT-MySQL would not stay up. "morpheus-ctl tail mysql" told me why:[ERROR] [MY-011971] [InnoDB] Tablespace 'innodb_undo_002' Page
[page id: space=4294967278, page number=41] log sequence number 52083774429
is in the future! Current system log sequence number 52081909523.
[ERROR] [MY-011972] [InnoDB] Your database may be corrupt or you may have copied
the InnoDB tablespace but not the InnoDB redo log files.
[ERROR] [MY-012153] [InnoDB] Trying to access page number 1751479667 in space
4294967278, space name innodb_undo_002, which is outside the tablespace bounds.
[ERROR] [MY-013183] [InnoDB] Assertion failure: fil0fil.cc:7529Three things worth pulling out of that:"log sequence number ... is in the future"The undo tablespace is holding pages written after the current redo position. Redo and undo are out of sync."outside the tablespace bounds"MySQL asked for a page that isn't in the file. Something got truncated or garbled."Assertion failure"InnoDB gave up on purpose. ROOT CAUSE ANALYSIS-There were two separate problems here, which took me a while to untangle because at firstthey looked like one.1) ENGINE-LEVEL CORRUPTION (GLOBAL)The "innodb_undo_002" tablespace held pages whose LSN was ahead of the redo log. The usual cause is an inconsistent VM snapshot or a partial restore. Roll back a snapshot that caught the data files but not the matching redo logs and this is roughly what you end up with. A storage-level hiccup does the same job.Page LSN: 52083774429Current Redo: 52081909523^ ~1.86 million LSN gap - missing redo records2) ROW-LEVEL DATA CORRUPTION (ISOLATED)Once I did get MySQL up, one row - id=73410 in "operation_data" - killed the engine every single time its "raw_data" LONGTEXT column was read. The bad bytes were sitting in the off-page LOB (Large Object) storage for that one field.The split matters here. The engine problem needed a fresh InnoDB. The row problem needed careful extraction. Fixing either one on its own would have left me where I started. WHY NOT JUST REPAIR IT IN PLACE?-Three reasons for this -A corrupt undo tablespace isn't something a table-level repair can touch. It's structural. The only clean fix is to reinitialise the InnoDB data directory so it regenerates undo and redo from scratch.Setting "raw_data" to NULL in place doesn't work either. That update makes InnoDB walk and free the corrupt LOB pages, which is precisely the read that crashes it. You'd be asking the engine to do the one thing it can't survive.And the appliance had to stay recoverable the entire time. One careless "rm" and the data is gone for good, with no second attempt So the order became: get everything out first, rebuild second. THE PLAN: DUMP, REBUILD, REIMPORT-Every step had to be reversible. The sequence I settled on:1. FORCE-START MySQL in recovery mode to get read access back.2. DUMP everything reachable to SQL files while running under forced recovery.3. ISOLATE the rows that crash the engine and rebuild them without the bad field.4. REBUILD the MySQL data directory from scratch, giving me a clean InnoDB with freshundo and redo logs.5. REIMPORT the recovered dumps into the new database.6. REPAIR the collateral damage (the Elasticsearch indices) and bring the appliance up.One rule I stuck to the whole way through: never "rm" the original corrupt data. Move it, copy it, rename it, but nothing gets deleted until the recovery is verified. STEP-BY-STEP RECOVERY-PHASE 1 - FORCE-START AND ASSESS"innodb_force_recovery" is a dial from 1 to 6. One is mild. Six is a last resort and generally means you're getting swords out of a lost battle. Start at the bottom and only climb if you have to. Level 2 was enough here.# /etc/morpheus/morpheus.rb
mysql['innodb_force_recovery'] = 2
sudo morpheus-ctl reconfigure
sudo morpheus-ctl start mysqlWith MySQL readable again, I went through the tables one at a time to work out what was actually salvageable.PHASE 2 - DUMP EVERYTHING RECOVERABLEEverything except the poisoned table first:/opt/morpheus/embedded/mysql/bin/mysqldump -h 127.0.0.1 -u root -p \
--quick --skip-lock-tables --no-tablespaces \
--ignore-table=morpheus.operation_data \
morpheus > /tmp/morpheus_no_operation_data.sqlThen "operation_data". A straight dump died around row 35,684, so I binary-searched the ID range, halving the failing window each time:70001-73411 -> failed
71701-73411 -> failed
72551-73411 -> failed
...
73304-73411 -> failed at row 106That got me down to a single row:SELECT id FROM operation_data WHERE id=73410; -- OK (row header intact)
SELECT * FROM operation_data WHERE id=73410; -- CRASHES MySQLThen column by column, until I found the one that was actually bad:SELECT raw_data FROM operation_data WHERE id=73410; -- CRASHES-- every other column returned fineFrom there I built a custom dump that rebuilt row 73410 with "raw_data" set to NULL and every other field intact.PHASE 3 - REBUILD THE DATA DIRECTORYWith the dumps safe on disk, time to rebuild.sudo morpheus-ctl stop morpheus-ui
sudo morpheus-ctl stop mysql
# Remove the force_recovery line from /etc/morpheus/morpheus.rb first
# Preserve the corrupt data - never delete it
sudo cp -a /var/opt/morpheus/mysql/data /var/opt/morpheus/mysql/data.corrupt.$(date +%F)
sudo mv /var/opt/morpheus/mysql/data /var/opt/morpheus/mysql/data.old
sudo morpheus-ctl reconfigure # regenerates a clean InnoDB with fresh undo/redoPHASE 4 - FIX THE COLLATERAL DAMAGE (ELASTICSEARCH)Whatever hit MySQL had also left Elasticsearch in a red cluster state, and that was blocking "reconfigure". These indices are only search mirrors of MySQL data, so dropping the broken ones is safe. Morpheus repopulates them on its own:for idx in $(curl -s 'http://localhost:9200/_cat/indices?h=health,index' \
| awk '$1=="red"{print $2}'); do
curl -s -X DELETE "http://localhost:9200/$idx"
donePHASE 5 - REIMPORT THE RECOVERED DATARecovered tables into the clean database first, then the repaired "operation_data":SOCK=/var/run/morpheus/mysqld/mysqld.sock
( echo "SET FOREIGN_KEY_CHECKS=0; SET UNIQUE_CHECKS=0;"; \
cat /root/morpheus_recover/morpheus_no_operation_data.sql; \
echo "SET FOREIGN_KEY_CHECKS=1;" ) \
| /opt/morpheus/embedded/mysql/bin/mysql -u root -p --socket=$SOCK morpheus
/opt/morpheus/embedded/mysql/bin/mysql --binary-mode -u root -p \
--socket=$SOCK morpheus < /root/morpheus_recover/operation_data_final.sqlPHASE 6 - BRING IT BACK UPsudo morpheus-ctl start morpheus-ui
sudo morpheus-ctl status WHERE IT WENT WRONG-None of that ran start to finish without stopping. Six things were what made this entire process long and painful for me.1) "reconfigure" HUNG AT "mysql_sleep"reconfigure waits for MySQL to answer, and MySQL was down. Start MySQL first, then run reconfigure. Obvious once you see it.2) "reconfigure" HUNG AT "elasticsearch_wait"The red indices again. Dropping them unblocked it.3) IMPORT FAILED - "TABLE DOESN'T EXIST"My recovered dump was INSERTs only, no CREATE TABLE, because I'd excluded the table from the main dump in the first place. I pulled the schema off a healthy Morpheus box.4) IMPORT FAILED - SYNTAX ERROR NEAR ",NULL)"This one took a while. "od -c" showed a raw 0x01 byte sitting where a 0 or 1 should be, which is the literal binary value of the "enabled" BIT(1) column. A perl substitution sorted it:perl -pe 's/,\x00,/,0,/g; s/,\x01,/,1,/g' input.sql > output.sql5) IMPORT FLOODED WITH "PAGER SET TO STDOUT"The dump had literal "\n" as text instead of actual newlines, so the mysql client read "\P" as the pager command and announced it, over and over. Convert the escapes and import with "--binary-mode":perl -pe 's/\\n/\n/g' input.sql > output.sql
mysql --binary-mode -u root -p morpheus < output.sql 6) GREP AND AWK GAVE ME WRONG NUMBERSEach INSERT was one 10KB+ line, which blows past line-length limits in both tools. They don't warn you, they just miscount. Slurping the file in perl gave honest answers:perl -0777 -ne '$c=()=/pattern/g; print "$c\n"' file.sql THE OUTCOME-Checks after the reimport:SELECT COUNT(*) FROM operation_data; -- 35687
SELECT COUNT(*) FROM operation_data WHERE enabled = 1; -- 35687
SELECT raw_data IS NULL FROM operation_data WHERE id=73410; -- 1
SELECT MAX(id) FROM operation_data; -- 73411  METRIC RESULTTotal Rows - 35,687Rows Recovered - 35,687 (100%)Fields Lost - 1 (raw_data at row id = 73410)Tables Recovered - AllElasticsearch Indices - Rebuilt from MySQL The appliance came back to full service. One field short of everything. WHAT I TOOK AWAY FROM IT-On prevention:* Back MySQL up automatically, then actually test the restore. An untested backup is a guess.* Settle MySQL before snapshotting the VM. Application-consistent snapshots avoid the exact redo/undo mismatch that started all this.* Read the InnoDB warnings. Corruption tends to mutter in the logs for a while before it shouts.On recovery:---> "mv", don't "rm". Nothing gets deleted until the new database is verified.---> Start "innodb_force_recovery" low. Every level up costs you something.---> Binary-search the bad rows, then narrow to the column. Testing 35,000 rows one at a time will eat your day.---> Check the dump files before you import them. Control bytes and escaping problems are common, and "--binary-mode" exists for a reason.WRAPPING UP-InnoDB corruption at this level looks terminal. Usually it isn't. Dump what you can reach, isolate what you can't, rebuild the engine clean, put the data back. None of these steps are clever on their own.What makes the difference is being willing to go slowly, and treating the corrupt files as untouchable until you've proved the new database works. Force-start, extract, isolate, rebuild, reimport, verify. In that order and without shortcuts.One field lost out of 35,687 rows. I'll always take that over a lost Morpheus appliance.

Схожие новости

#Наименование новостиТональностьИнформативностьДата публикации
1Каждые 5 минут транзакции в PostgreSQL замирают на 3-7 секунд. ...011.4328-09-2026
2Issues with fiber channel multipathing on HPE ProLiant DL385?08.7827-09-2026
3База данных захлебывается в дисковом I/O wait, хотя на сервере ...08.4126-09-2026
4Каждые 5 минут транзакции в PostgreSQL замирают на 3 - ...-111.8628-09-2026
5База данных захлебывается в дисковом I/O wait, хотя на сервере ...08.626-09-2026
6HP DL380p Gen8 – Replacement drive immediately showing Failed013.9725-09-2026
7Opera GX 136.0.6008.50 crashing repeatedly on startup – possible Chromium&#x2F;DirectWrite font issue017.3624-09-2026
8Каждые пять минут нагруженный PostgreSQL кластер словно падает в обморок ...08.0327-09-2026
9Re: HPE Nimble storage controller HA issue08.8126-09-2026
10Re: CRITICAL: Prolonged Cluster Partition / Split-Brain on 3-Node HPE VME 8 Cluster010.4127-09-2026

Классификация: . Схожих патентов: 0. Схожих новостей: 10. Тональность: 0. Информативность: 7.64. Источник: community.hpe.com.