Just like our own skills, Claude skills and the context in them can quietly degrade over time without updates. Anthropic itself recommends updating skills rather than leaving them as finished markdown files. Think of them as complete toolkits or mini-experts who keep on improving with time and experience.
I'm still getting a handle on using skills for workflow automation properly. Continuously patching a skill whenever something goes wrong seems tiresome and the start of bloated instructions, even though Claude skills supports versions. So I'm trying a workaround. I'm using NotebookLM as a more objective "maintenance guy" for the Claude skills.
I started with a simple meal skill
Even a mundane task can change quickly
My test skill creates a 21-day meal plan based on my preferences (though I have simulated the dishes here), available ingredients, cooking time, and leftovers. It also produces a shopping list so I don't have to work out what to buy separately. The instructions tell Claude to avoid repeating the same main dish too often, reuse ingredients across meals, account for leftovers, and keep the shopping list practical.
Maybe, a bit simple and silly for an AI experiment. But that simplicity allows me to see how my lifestyle and even random events in a single week can throw off the plan in a SKILLS.md file. It's always possible to tweak and scale it up for more complex Claude skills.
The critical part is that meal planning has so many variables. An ingredient may not be in stock. A busy week can throw off the entire plan Claude makes for me. Even a day out can have a cascading effect on the meal prep. So, the kitchen isn't a bad Petri dish for the Claude skills maintenance test.
One week can go wrong, while the next runs perfectly. So it's better to have 3–4 weeks minimum of data before running the skill.
I log corrections instead of patching
Real-world use exposes the missing rules
Rather than immediately changing the skill whenever something doesn't work, I keep a simple Google Doc recording what happened during each week's meal planning. As I'm using NotebookLM for this test, I record the log in Google Docs, and not as an uploaded markdown file. NotebookLM syncs Drive sources periodically, so it can self-update the source. It's not always an instant real-time sync. Often, you might have to manually open the NotebookLM "Sources" panel and click the Refresh icon on the Google Doc source to force it to pull the latest entries.
I enter things such as meals I skipped, ingredients I couldn't find, recipes that took too long, leftovers that weren't used, and suggestions I had to correct. I also record things that worked particularly well. For example, the skill might schedule a lentil curry on Monday and suggest using the leftovers in a pancake for lunch on Tuesday. But if I discover that the leftovers aren't enough for two people, that's useful information about how the skill should plan portions.
The critical mindset shift here is to not create or change rules with every miss. Simply recording them as observations (as scientists do in actual lab experiments) is an exercise in finding patterns. One meal gone wrong doesn't mean inserting another rule into the Claude skill.
Editing the SKILL.md file is fine for fixing genuine mistakes like typos or a rule you forgot to include. Just highlight text in a skill file and click Edit.
I let NotebookLM find the patterns
It compares my rules with actual experience
After several weeks, I upload the current SKILL.md and my meal-planning log to a NotebookLM notebook. NotebookLM can then compare and analyze the skill alongside my actual observations.
I then ask NotebookLM to compare what the skill says with what really happened. Here's the exact prompt I use:
Act as a maintenance reviewer for my Claude weekly meal-planning skill.
I have provided two sources:
1. The current SKILL.md file, which describes how Claude is supposed to create my weekly meal plans.
2. A running meal-planning log containing what actually happened when I used the Skill.
Compare the Skill file against the meal-planning log. List every place where what happened doesn't match what the skill says. Cite the log entry for each gap and sort by how often it happened.
Classify every recommendation as:
PROMOTE — Strong evidence suggests the Skill should be changed.
WATCH — Interesting pattern, but there isn't enough evidence to change the Skill yet.
REJECT — This is probably a one-off incident or does not justify changing the Skill.
Important rules:
- Do not rewrite the Skill.
- Do not invent preferences that aren't supported by the meal-planning log.
- Do not turn a single disliked meal into a permanent rule.
- Look for patterns across multiple weeks.
- Prefer small, precise changes over adding lots of new instructions.
- Distinguish between a genuine problem with the Skill and a one-off problem caused by unusual circumstances.
- Pay attention to successful meals too. Repeated successes may reveal useful preferences that aren't currently encoded in the Skill.
End with a section called "Most important changes" containing only the strongest PROMOTE recommendations.
This is the part I like most about the workflow. The skill is the baseline for my ideal system, while the meal log catches my actual behavior across the weeks. NotebookLM helps me see the difference between the two. I could have skipped NotebookLM by dropping both files into a Claude Project and asking Claude to review its own skill.
That works, but a fresh reviewer works better. Anthropic's own guidance splits the job between one Claude that refines a skill and another that uses it. NotebookLM answers only from your sources, and you can trace everything to its real log entry.
I designed the prompt with Claude's help. You can also let Claude help you fine-tune the prompt and make it bulletproof for your scenario.
I turn patterns into recommendations
Not every meal mistake needs a new rule
Once NotebookLM identifies recurring problems, I ask it to turn the strongest findings into a "changelog." NotebookLM might suggest making the rule more specific. Thanks to NotebookLM's grounded sources, the responses come from my real-world use rather than from a model's or my own assumptions about what sounds like a good meal-planning rule.
For instance, it can suggest reusing ingredients where practical, but avoid requiring unusually large quantities of perishable ingredients merely to reduce the number of items on the shopping list. But I have also made the mistake of fixing a rule that broke another rule. That led to errors in Claude's response.
That's why the Claude prompt recommended the Promote, Watch, and Reject categories. This prevents the skill from growing every time something goes slightly wrong. Otherwise, maintaining it could eventually become an exercise in over-fixing by adding exceptions on top of exceptions.
I test the updated skill again
The new rule has to survive the future weeks
After reviewing the proposed changes, I open the skill in a Claude chat and apply the changes. I pay particular attention to the meals and situations that forced the changes. I check whether the next meal plan actually solves that problem without taking the rest of the plan off-track.
It's a regular maintenance cycle. Maybe it's overkill for a simple exercise like meal planning, but it's critical for more important tasks. If an update backfires, your changelog is the only record of what changed.
The process also makes it easy to roll back anything that doesn't work. I can look at the correction log and see the real-world problems that caused it.
Pick one skill you use the most
Try this with a Claude skill you already use regularly. Don't spend hours trying to make the initial version perfect. Use it, record what goes wrong, and let several real-world examples accumulate. For instance, you can keep a running log of corporate jargon or stylistic habits you keep having to manually edit out. Coders can log syntax errors generated by Claude.
Then put the skill and your log into NotebookLM and ask it to find the patterns. Now, a Claude skill isn't a sacrosanct set of instructions. You're treating it as a small system that gets better as you learn what actually works with your behaviors.



















