File structure cleanup
Audits a workspace file inventory for duplicate names, stale versions, misplaced files, inconsistent naming, and empty folders, then proposes a deterministic cleanup map without deleting or moving anything. Use for shared-drive housekeeping, repository organization, and document-library maintenance.
Veröffentlicht 21. Aug. 2026 · Aktualisiert 26. Aug. 2026
Voraussetzungen
Map inventory columns, protected paths, document classes, and folder ownership. Add naming, retention, and approved-canonical rules. Define the confidence labels for exact duplicate, likely duplicate, stale version, and naming defect.
Skill-Dokument
Das vollständige SKILL.md, das dein Agent liest und befolgt.
File structure cleanup
Purpose
Turn a file inventory into a safe cleanup plan. The plan identifies exact paths, duplicate groups, canonical candidates, naming defects, and proposed folder destinations while preserving evidence for every recommendation.
Scope
Assess file paths, filenames, extensions, sizes, modified dates, owners, folder names, and version markers in the supplied inventory. Apply the company's naming and retention rules where available.
Excluded: deleting, moving, renaming, restoring, or changing permissions on files.
Data basis
- Workspace file inventory with full path, filename, extension, size, modified timestamp, owner, and checksum where available.
- Folder tree export and naming, retention, or document-classification policy.
- Optional protected-path list and known canonical-file register.
Result
A cleanup register with issue type, affected path, duplicate group, canonical candidate, proposed destination/name, evidence, and confidence, plus a summary of counts and storage impact.
Quality criteria
- No path is recommended for deletion without a duplicate or retention-rule citation.
- Checksum matches outrank filename similarity when identifying duplicate content.
- Protected paths and current files are excluded from destructive recommendations.
- Every proposed move or rename retains the original path.
Instructions
Treat full path and checksum as identifiers. Group exact duplicates before near-duplicates. A newer modified date alone does not establish the canonical file when a protected or approved register exists. Preserve extensions and distinguish drafts, signed versions, and superseded copies. Use the supplied retention period, not an invented age threshold. Mark inaccessible metadata as unassessed and never infer ownership from folder names.
Adapt before use
- Map inventory columns, protected paths, document classes, and folder ownership.
- Add naming, retention, and approved-canonical rules.
- Define the confidence labels for exact duplicate, likely duplicate, stale version, and naming defect.
Process detail
Inventory paths and metadata
Count files and folders, normalize path separators, and check for missing extensions, timestamps, checksums, and owners.
Data basis: File inventory and folder tree columns.
Result: A stable inventory keyed by full path.
Acceptance criterion: File and folder counts match the source export and every row has a path.
Exception: Keep inaccessible or malformed rows in an exception set.
Group exact duplicates
Group files by checksum and size, then list every path in each group.
Data basis: Checksum, byte size, extension, full path, and protected-path list.
Result: Exact-duplicate groups with candidate canonical paths.
Acceptance criterion: Each file belongs to at most one exact-duplicate group and no protected path is proposed for deletion.
Exception: If checksum is absent, classify the group as similarity-only rather than exact.
Detect likely versions and naming defects
Compare stem names, version markers, dates, extensions, and folder context to find near-duplicates, stale copies, and naming violations.
Data basis: Filename, parent folder, modified date, owner, document class, and naming policy.
Result: Issue rows for version clusters and naming defects.
Acceptance criterion: Each issue includes the matched paths and the specific naming or version rule breached.
Exception: Do not merge files with different extensions or document classes without policy evidence.
Choose a cleanup disposition
Assign keep, archive review, rename proposal, move proposal, or no action and identify a canonical candidate.
Data basis: Duplicate groups, protected paths, retention rules, and canonical-file register.
Result: Disposition map preserving original path and rationale.
Acceptance criterion: Every issue has one disposition, confidence, and cited rule or inventory evidence.
Exception: Use archive review when age exceeds policy but business retention cannot be established.
Quantify cleanup impact
Calculate duplicate count, potentially reclaimable bytes, stale-version count, and naming-defect count by folder.
Data basis: Disposition map, file sizes, and folder hierarchy.
Result: Cleanup summary with counts and byte totals.
Acceptance criterion: Reclaimable bytes include only exact duplicates not protected or canonical; totals tie to rows.
Exception: Keep byte totals separate by storage unit and state unknown sizes.
Document what could not be assessed
List files without checksum, inaccessible metadata, protected paths, ambiguous versions, and absent retention rules.
Data basis: Inventory exceptions and disposition map.
Result: Open-points section with path and reason.
Acceptance criterion: Every excluded path is cited and the summary count equals the exception register.
Exception: State zero exceptions only after checking all inventory rows.
Verwandte Skills
- Sync a spreadsheet into a record type
Loads an uploaded spreadsheet into a record type of the workspace with the bulk sync tool, previewing the changes before writing. Use for master data imports, periodic refreshes of a table, or migrating a list from Excel into records.
- Anomaly variance explanation
Explains material changes in a business metric by decomposing period, volume, price, mix, and data-quality effects, then ranks evidence-backed causes in a cited analysis sheet and management memo. Use for KPI variance reviews, monthly business reviews, forecast misses, anomaly investigation, and driver analysis.
- Data migration field mapping
Maps source-system fields to a target data model for migration, documenting transformations, requiredness, value translations, identifiers, validation rules, and unresolved collisions. Use for CRM migrations, ERP imports, warehouse loads, application replacement, and controlled spreadsheet-to-system field mapping.
- Dataset Quality Assessment
Dataset Quality Assessment converts table schema, column types, null policy, primary and foreign keys into profiling report and defect register with field metrics, duplicate groups, invalid values, orphan keys, temporal gaps, and severity, with cited evidence, explicit rules, and unresolved exceptions. Use for dataset quality assessment, recurring review, and decision preparation.