%20para%20n%C3%A3o%20levar%20em%20conta%20uma%20nova%20linha%20quando%20ela%20estiver%20entre%20aspas%20duplas%3F.png)
Eu tenho um csv
arquivo grande que começa assim:
codeRegion,nomEPCI,codeDepartement,nomCommune,populationTotale,titre,objet,site_web,libelle_objet_social1,libelle_objet_social2,siret,numero_waldec,nombreAnneesExistence,date_creation,nombreAnneesDerniereDeclaration,date_derniere_declaration,position_activite,date_dissolution,codeCommune,adresse_siege_complement,adresse_siege_numero_voie,adresse_siege_type_voie,adresse_siege_libelle_voie,adresse_siege_distribution,adresse_siege_code_postal,nom_declarant,adresse_gestion_complement_association,adresse_gestion_complement_geo,adresse_gestion_libelle_voie,adresse_gestion_distribution_facturation,adresse_gestion_code_postal,adresse_gestion_achemine,adresse_gestion_pays,civilite_declarant,codeEPCI
01,CA CAP Excellence,971,Abymes,54049,ABSOLU MAS/KA,"organiser des manifestations sociales et culturelles ainsi que des activités de loisirs","",Action socio-culturelle,"Clubs de loisirs, relations","",W9G2015492,0.5808219178082191,2020-02-05,0.5808219178082191,2020-02-05,A,0001-01-01,97101,62 Résidence les Lavandes,62,RES,Boisripeaux,_Boisripeaux,97139,"","",RÉSIDENCE LES LAVANDES,62 RESIDENCE BOISRIPEAUX,"",97139,LES ABYMES,FRANCE,PM,200018653
01,CA CAP Excellence,971,Abymes,54049,AMBITION PLUS DE TERRASSON,proposer des activités sportives et culturelles,"",Action socio-culturelle,"Sports, activités de plein air","",W9G2015457,0.6356164383561644,2020-01-16,0.6356164383561644,2020-01-16,A,0001-01-01,97101,Maison Andreopa Marcel,3,CHEM1,"Route de Terrasson, Rue Albert Léogane",_97139,97139,"","",MAISON ANDREOPA MARCEL,3 CHEMIN ROUTE DE TERRASSO,"",97139,LES ABYMES,FRANCE,PF,200018653
01,CA CAP Excellence,971,Abymes,54049,ASSOCIATION SYNDICALE DE LA RESIDENCE MORNE CARUEL,gérer et d'entretenir les espaces communs cette mission ne comporte pas la possibilité daliéner les espaces indivis si ce n'est au profit de la commune,"",Actions de sensibilisation et d'éducation à l'environnement et au développement durable,"","",W9G2015446,14.797260273972602,2005-11-21,14.797260273972602,2005-11-21,A,0001-01-01,97101,residence morne caruel,"","","",_,97139,"",residence morne caruel,"","","",97139,Les Abymes,FRANCE,PM,200018653
01,CA CAP Excellence,971,Abymes,54049,ASSOCIATION SPORTIVE DU LYCEE POLYVALENT CHEVALIER DE SAINT-GEORGES,"organiser et développer en prolongement de l'éducation physique et sportive donnée pendant les heures de scolarité, l'initiation et la pratique sportive pour les élèves qui y adhèrent elle représente l'établissement dans les épreuves sportives scolaires","",Activités de plein air (dont saut à l'élastique),"Centres de loisirs, clubs de loisirs multiples","",W9G2002394,19.87123287671233,2000-10-26,4.742465753424658,2015-12-09,A,0001-01-01,97101,"","",BD,des Héros,_Baimbridge,97139,"","",LYCéE POLYVALENT CHEVALIER SAINT-GEORG,BOULEVARD DES HéROS,BAIMBRIDGE,97139,ABYMES,FRANCE,PM,200018653
01,CA CAP Excellence,971,Abymes,54049,"LOISIRS, COMPETITIONS, CLUB (LO. CO. CLUB)","participer aux competitions cyclistes
organiser des manifesta- tions sportives a caractere promotionnel.","",Activités de plein air (dont saut à l'élastique),"Centres de loisirs, clubs de loisirs multiples","",W9G2010818,28.747945205479454,1991-12-13,27.378082191780823,1993-04-26,A,0001-01-01,97101,"","","","",_,97139,"","","","","",97139,LES ABYMES,FRANCE,PM,200018653
01,CA CAP Excellence,971,Abymes,54049,ASSOCIATION ETUDIANTE DE TOURISME ET DE LOISIRS.( A.T.O.L.),réunir les anciens élèves ayant eu une formation de tourisme et de loisirs et de promouvoir les actions du brevet de technicien supérieur de tourisme et de loisir,"","Amicales, personnel d’établissements scolaires ou universitaires","Syndicats d'initiative, offices de tourisme, salons du tourisme","",W9G2011048,24.53698630136986,1996-02-27,24.53150684931507,1996-02-29,A,0001-01-01,97101,"","","",Ecole superieure des cadres et techniciens,_Route de la rocade grand-camp,97139,"","","",Ecole superieure des cadres et technic,Route de la rocade grand-camp,97139,LES ABYMES,FRANCE,PM,200018653
Sua quinta linha de dados possui uma descrição, entre aspas duplas, contendo uma nova linha. Funciona perfeitamente comExcelouLibreCalc.
Preciso reter apenas as linhas que começam em uma região francesa específica. Por exemplo aqui, o '01' (Guadalupe).
Eu executo:
# Get the CSV header
head -n1 associations_touristiques.csv > associations_touristiques_gua.csv
# Extract data of region '01'
cat associations_touristiques.csv | grep -a '^01,' >> associations_touristiques_gua.csv
Mas minha final csv
está quebrada.
codeRegion,nomEPCI,codeDepartement,nomCommune,populationTotale,titre,objet,site_web,libelle_objet_social1,libelle_objet_social2,siret,numero_waldec,nombreAnneesExistence,date_creation,nombreAnneesDerniereDeclaration,date_derniere_declaration,position_activite,date_dissolution,codeCommune,adresse_siege_complement,adresse_siege_numero_voie,adresse_siege_type_voie,adresse_siege_libelle_voie,adresse_siege_distribution,adresse_siege_code_postal,nom_declarant,adresse_gestion_complement_association,adresse_gestion_complement_geo,adresse_gestion_libelle_voie,adresse_gestion_distribution_facturation,adresse_gestion_code_postal,adresse_gestion_achemine,adresse_gestion_pays,civilite_declarant,codeEPCI
01,CA CAP Excellence,971,Abymes,54049,ABSOLU MAS/KA,"organiser des manifestations sociales et culturelles ainsi que des activités de loisirs","",Action socio-culturelle,"Clubs de loisirs, relations","",W9G2015492,0.5808219178082191,2020-02-05,0.5808219178082191,2020-02-05,A,0001-01-01,97101,62 Résidence les Lavandes,62,RES,Boisripeaux,_Boisripeaux,97139,"","",RÉSIDENCE LES LAVANDES,62 RESIDENCE BOISRIPEAUX,"",97139,LES ABYMES,FRANCE,PM,200018653
01,CA CAP Excellence,971,Abymes,54049,AMBITION PLUS DE TERRASSON,proposer des activités sportives et culturelles,"",Action socio-culturelle,"Sports, activités de plein air","",W9G2015457,0.6356164383561644,2020-01-16,0.6356164383561644,2020-01-16,A,0001-01-01,97101,Maison Andreopa Marcel,3,CHEM1,"Route de Terrasson, Rue Albert Léogane",_97139,97139,"","",MAISON ANDREOPA MARCEL,3 CHEMIN ROUTE DE TERRASSO,"",97139,LES ABYMES,FRANCE,PF,200018653
01,CA CAP Excellence,971,Abymes,54049,ASSOCIATION SYNDICALE DE LA RESIDENCE MORNE CARUEL,gérer et d'entretenir les espaces communs cette mission ne comporte pas la possibilité daliéner les espaces indivis si ce n'est au profit de la commune,"",Actions de sensibilisation et d'éducation à l'environnement et au développement durable,"","",W9G2015446,14.797260273972602,2005-11-21,14.797260273972602,2005-11-21,A,0001-01-01,97101,residence morne caruel,"","","",_,97139,"",residence morne caruel,"","","",97139,Les Abymes,FRANCE,PM,200018653
01,CA CAP Excellence,971,Abymes,54049,ASSOCIATION SPORTIVE DU LYCEE POLYVALENT CHEVALIER DE SAINT-GEORGES,"organiser et développer en prolongement de l'éducation physique et sportive donnée pendant les heures de scolarité, l'initiation et la pratique sportive pour les élèves qui y adhèrent elle représente l'établissement dans les épreuves sportives scolaires","",Activités de plein air (dont saut à l'élastique),"Centres de loisirs, clubs de loisirs multiples","",W9G2002394,19.87123287671233,2000-10-26,4.742465753424658,2015-12-09,A,0001-01-01,97101,"","",BD,des Héros,_Baimbridge,97139,"","",LYCéE POLYVALENT CHEVALIER SAINT-GEORG,BOULEVARD DES HéROS,BAIMBRIDGE,97139,ABYMES,FRANCE,PM,200018653
01,CA CAP Excellence,971,Abymes,54049,"LOISIRS, COMPETITIONS, CLUB (LO. CO. CLUB)","participer aux competitions cyclistes
01,CA CAP Excellence,971,Abymes,54049,ASSOCIATION ETUDIANTE DE TOURISME ET DE LOISIRS.( A.T.O.L.),réunir les anciens élèves ayant eu une formation de tourisme et de loisirs et de promouvoir les actions du brevet de technicien supérieur de tourisme et de loisir,"","Amicales, personnel d’établissements scolaires ou universitaires","Syndicats d'initiative, offices de tourisme, salons du tourisme","",W9G2011048,24.53698630136986,1996-02-27,24.53150684931507,1996-02-29,A,0001-01-01,97101,"","","",Ecole superieure des cadres et techniciens,_Route de la rocade grand-camp,97139,"","","",Ecole superieure des cadres et technic,Route de la rocade grand-camp,97139,LES ABYMES,FRANCE,PM,200018653
01,CA CAP Excellence,971,Abymes,54049,ASSOCIATION AMICALE DES AGENTS D'ENTRETIEN DE LA MUNICIPALITE DES ABYMES.,"realiser des rencontres et echanges d'ordre culturel, sportif, social avec toutes categories de personnel communal et autres associations ou groupements.","",Association du personnel d'une entreprise (hors caractère syndical),"Comités de défense et d'animation de quartier, association locale ou municipale","",W9G2012061,31.56986301369863,1989-02-16,31.56986301369863,1989-02-16,A,0001-01-01,97101,"","","","",_,97139,"","","","","",97139,LES ABYMES,FRANCE,PM,200018653
A linha 5 é truncada apósciclistas.
Qual a maneira correta de fazer cat
o comando não levar em conta uma nova linha quando está entre aspas duplas,
e é necessário alterar o grep
comando também depois disso, e em caso afirmativo: como?
Responder1
Usando csvgrep
docsvkit
pacote para extrair todos os registros que possuem um codeRegion
valor contendo a string 01
:
csvgrep -c codeRegion -m 01 file.csv
Isso está usando um analisador CSV adequado, portanto, não haverá problemas com novas linhas ou vírgulas em campos citados corretamente.
A -c
opção seleciona a coluna que gostaríamos de investigar, por número ou por nome, e -m
designa a string com a qual combinar. Pode-se também usar -r
para combinar com uma expressão regular, por exemplo, -r '^01$'
para evitar a correspondência de strings onde 01
há uma substring (como em 011
). Ver csvgrep --help
.
Responder2
awk '/^01/||n%2{print;n+=gsub(/"/,"&")}' file
Para cada linha,
/^01/||n%2
Se a linha começar com01
oun
(inicialmente zero) for ímpar,print
Impriman+=gsub(/"/,"&")
incrementarn
pelo valor de retorno dagsub
função.
Isso substitui todas as aspas duplas/"/
por si mesmo"&"
. Isso seria inútil, de fato, mas também retorna o número de substituições feitas, portanto é uma forma de contar o número de aspas duplas na linha.
Observe que se n
for ímpar ( n%2
), a linha não tem aspas duplas de fechamento, então ela continua imprimindo até n
ser par, independentemente de haver uma /^01/
correspondência nas próximas linhas.
Uma comparação lado a lado para você:
$ diff -yW 30 <(cat file) <(awk '/^01/||n%2{print;n+=gsub(/"/,"&")}' file)
04,xde <
01,abc" 01,abc"
cd cd
as" as"
02,dsad <
03,1ad" <
01,as,"as 01,as,"as
us" us"
02,s <
01,a 01,a